Haiyang Zhou

dblp:76/6502 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0004-3616-9120ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 360Explorer: Exploring 4D Controllable World in Panoramic Videos
abstract
We present 360Explorer, a novel approach for generating 4D controllable panoramic videos conditioned on user-provided 3D instructions for exploring and manipulating dynamic worlds. Compared to existing perspective-based methods struggle to address spatial consistency during camera rotation in place, we introduce the panoramic view in controllable video generation models to inherently maintain the view recall consistency. By introducing dynamic point clouds as the 4D scene representations, 360Explorer unifies the modeling of camera transformations and object movements as incomplete renders to describe precise control instructions in 3D worlds. To tackle the data limitation in acquiring multi-viewpoint panoramic videos, we further propose a reverse warping strategy to construct the training dataset on easily accessible monocular panoramic videos. Extensive experiments demonstrate that 360Explorer achieves superior performance in creating 4D controllable panoramic videos with camera transformation and object movements aligned with diverse provided instructions.
Xinhua Cheng, Haiyang Zhou, Wangbo Yu, Tanghui Jia, Bin Lin 0014, Yunyang Ge, Li Yuan 0007
AAAI2
2026 HoloDreamer: Holistic 3D Panoramic Scene Generation From Text Descriptions
abstract
3D scene generation is in high demand across various domains, including virtual reality, gaming, and the film industry. Owing to the powerful generative capabilities of text-to-image diffusion models that provide reliable priors, creating 3D scenes using only text prompts has become viable, thereby significantly advancing research in text-driven 3D scene generation. Prevailing methods typically employ the diffusion model to generate an initial local image, followed by iteratively outpainting the local image to gradually generate scenes. Nevertheless, these outpainting-based approaches are prone to producing globally inconsistent results with low completeness, restricting their broader applications. To tackle these problems, we introduce HoloDreamer, a framework that begins by generating a high-definition panorama to holistically initialize the full scene, and leverages 3D Gaussian Splatting (3D-GS) for rapid 3D scene reconstruction, thereby facilitating the creation of view-consistent and fully enclosed 3D scenes. Specifically, we propose Stylized Equirectangular Panorama Generation, a pipeline that combines multiple diffusion models to enable stylized and detailed equirectangular panorama generation from complex text prompts. Subsequently, Enhanced Two-Stage Panorama Reconstruction is introduced, conducting a two-stage optimization of 3D-GS to inpaint the missing region and enhance the integrity of the scene. Comprehensive experiments demonstrated that our method outperforms prior works in terms of overall visual consistency and harmony, as well as reconstruction quality and rendering robustness when generating fully enclosed scenes.
Haiyang Zhou, Xinhua Cheng, Wangbo Yu, Yonghong Tian 0001, Li Yuan 0007
IEEE Trans. Vis. Comput. Graph.1
2025 An Energy-Efficient High-Utilization Hardware Architecture for Attention Mechanism in Transformer using Balanced Systolic Array and Multi-Row Interleaved Operation Ordering
abstract
Transformer-based neural networks have achieved remarkable performance. Designing energy-efficient and high-speed accelerators for the attention mechanism, which dominates the energy and latency in Transformers, has become increasingly significant. Existing attention accelerators commonly use algorithm-hardware co-design to achieve higher energy efficiency and speed. However, deeply customized algorithms make these accelerators dependent on a particular application. Therefore, optimizing hardware architecture is crucial for achieving general-purpose acceleration. We observe two limitations in the hardware architecture of existing attention accelerators. First, the widely used input stationary, weight stationary, and output stationary systolic arrays (SAs) can’t balance data reuse, register saving, and utilization, which hinders to build more energy-efficient and faster SA-based accelerators. Second, layer-by-layer operation ordering introduces high SRAM access overhead of intermediate results. To address the first limitation, we propose the “Balanced Systolic Array”, which improves energy efficiency by 40% compared to conventional systolic arrays and achieves a utilization rate of 99.5%. To address the second limitation, we propose “Multi-Row Interleaved” operation ordering, which reduces the SRAM energy by 31.7% By integrating two techniques, the proposed attention accelerator achieves a 39% improvement in energy efficiency and a 38% enhancement in throughput×energy efficiency compared to previous works.
Haiyang Zhou, Hongyang Hu, Jinshan Yue, Hanghang Gao, Yuanlu Xie, Xiaoxin Xu, Chunmeng Dou, Ming Liu 0022
DAC1
2025 Volatility Prediction and Classification Using Garch Model and Dynamic Random Forest With Gaussian Mixture Model
abstract
This study integrates traditional statistical models with advanced predictive techniques to analyze volatility in financial markets. The Generalized Autoregressive Conditional Heteroskedasticity (GARCH) model is employed to capture the autoregressive conditional heteroskedasticity in financial data, improving upon the constant variance assumption of earlier economic models. However, recognizing the limitations of GARCH in reflecting the asymmetric impact of positive and negative shocks, we incorporate models like GJR-GARCH and EGARCH that account for the leverage effect. To further enhance prediction accuracy and capture complex market dynamics, we leverage the Dynamic Random Forest (DRF) model. This ensemble learning method considers multiple factors simultaneously and updates volatility data dynamically through sliding windows. We cluster the predicted volatilities into high and low volatility periods using the Gaussian Mixture Model (GMM). Our numerical analysis demonstrates that the DRF model achieves high prediction accuracy, and the clustering approach effectively distinguishes between high and low volatility states, with significant differences in Value at Risk (VaR) and Tail Value at Risk (TVaR) between the two states. This methodology provides valuable insights for risk management and investment strategy formulation.
Haiyang Zhou
HPCC1
2025 HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
Haiyang Zhou, Wangbo Yu, Jiawen Guan, Xinhua Cheng, Yonghong Tian 0001, Li Yuan 0007
ACM Multimedia1
2024 A 2T P-Channel Logic Flash Cell for Reconfigurable Interconnection in Chiplet-Based Computing-In-Memory Accelerators
abstract
In this work, we propose a two-transistor (2T) p-type channel (p-channel) logic-compatible flash cell. Compared to the previous designs, the proposed structure features reduced area-cost and enhanced ability to pass through the logic ‘1’. Due to these advantages, we explore its application as the reconfigurable interconnections in the chiplet-based system. By integrating them into the silicon interposer, the 2T p-channel flash cells can potentially lead to the dense and flexible interconnection between multiple computing-in-memory (CIM) chiplets, resulting in highly reconfigurable and scalable chiplet-based CIM accelerators. A 180nm 1Kb 2T p-channel flash cell array is fabricated and characterized. The characterization results show the 2T p-channel flash cells exhibit a signal ratio >103over 1000 program/erase (P/E) cycles and the device-to-device variations are less than 21.07%. Their typical behaviors as routers are also confirmed by circuit simulations.
Weizeng Li, Linfang Wang, Zhi Li 0062, Wang Ye, Zhidao Zhou, Haiyang Zhou, Hanghang Gao, Jinshan Yue, Hongyang Hu, Fengman Liu, Chunmeng Dou
ISCAS6
2024 Prompt2Poster: Automatically Artistic Chinese Poster Creation from Prompt Only
abstract
As a critical component in graphic design, artistic posters are widely applied in the advertising and entertainment industry, thus the automatic poster creation from user-provided prompts has become increasingly desired recently. Although existing Text2Image methods create impressive images aligned with given prompts, they fail to generate ideal artistic posters, especially with Chinese texts. To create desired artistic Chinese posters including an aligned background, reasonable layouts, and stylized graphical texts from given prompts only, we propose an automatic poster creation framework, named Prompt2Poster. Our framework utilizes the capacity of the powerful Large Language Model (LLM) to extract user intention from provided prompts and generate the aligned background. Although only taking a user prompt as the input, linguistic, visual, and geometrical information is fully utilized in the framework, bringing the ability to fit different distributions. To achieve the use of multi-modal information in the framework, two carefully designed modules, Controllable Layout Generator (CLG) and Graphical Text Generator (GTG) are proposed, leading to accurate and pleasurable visual results. Comprehensive experiments demonstrate that our Prompt2Poster achieves superior performance, especially in text quality and visual harmony.
Yunyang Ge, Liuhan Chen, Haiyang Zhou, Qian Wang 0062, Xinhua Cheng, Li Yuan 0007
ACM Multimedia4
2023 A 40-nm SONOS Digital CIM Using Simplified LUT Multiplier and Continuous Sample-Hold Sense Amplifier for AI Edge Inference
abstract
Digital computing in memory (CIM) exhibits high precision as well as high energy efficiency (EE) yet still lacks discussion in nonvolatile memory (NVM). In this article, we propose a 40-nm silicon-oxide-nitride-oxide-silicon (SONOS)-based digital NVM CIM macro (DNV-CIM) featuring: 1) a simplified lookup table multiplier (SLUTM) combined with a lookup table (LUT) mapping scheme to improve area and EE and 2) a continuous sample-hold sense amplifier (CSH-SA) with an optimized voltage clamper and comparator for continuous read to reduce overall energy and time consumption for deep neural network (DNN) inference tasks. Performance evaluations indicate that the proposed DNV-CIM can achieve 93.04% accuracy and an EE up to 39.9 TOPS/W when running a 4-bit quantized ResNet18 trained on the CIFAR-10 dataset. This work presents a highly efficient digital CIM solution that can be readily implemented with commodity NVM.
Hongyang Hu, Haiyang Zhou, Danian Dong, Jinshan Yue, Wan Pang, Xiaoxin Xu, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.4
2022 Research on the selection of charging stations by Q-learning optimized AHP
abstract
The selection of charging stations is often accompanied by the mutual constraints of multiple criteria, such as driving time, waiting time, and deviation coefficient. Analytic hierarchy process (AHP) is widely used to solve complex multi-objective problems. In this paper, the Q-learning optimized AHP is proposed to solve the problem of charging station selection. Compared with the original AHP, the scientific and objective weight of AHP has been improved which makes the selection result of charging station more reliable. The simulation results show that the multi-criteria method proposed can effectively reduce the average cost of driving time and waiting time, compared with the method that only considers the total driving time, charging waiting time and deviation coefficient.
Tong Wang 0005, Haiyang Zhou
VTC Fall2
2018 Applying rotation-invariant star descriptor to deep-sky image registration
Haiyang Zhou, Yunzhi Yu
Frontiers Comput. Sci.1