Shihao Han

dblp:155/1098 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic Data
abstract
Precise and controllable image editing, especially object removal and insertion, represents one of the most common demands in image manipulation. However, existing methods suffer from severe limitations. Mask-based inpainting often introduces visual artifacts and semantic inconsistencies, while instruction-based approaches lack accurate spatial control and tend to unintentionally modify background regions. To address these issues, we propose two key contributions. First, we develop a fully automated and self-improving pipeline for synthetic data generation. This pipeline utilizes a Large Language Model (LLM) to generate diverse prompts, a Diffusion Transformer (DiT) fine-tuned evolutionarily to synthesize high-quality images, and a Multimodal LLM (MLLM) combined with open-set object detector for automated quality control and annotation. This process produces the Remove/Add Dataset (RAD), consisting of over 514,510 high-quality image pairs, each richly annotated with bounding boxes, segmentation masks, and a variety of editing instructions. Second, based on RAD, we introduce Remove/Add Anything (RAA), a novel editing framework with precise spatial control. Built upon a diffusion-based inpainting model, RAA achieves high editing accuracy by conditioning on both textual instructions and an explicitly defined region of interest (ROI), enabling efficient fine-tuning while maintaining global visual coherence. Extensive experiments demonstrate that RAA significantly outperforms existing open-source methods on both addition and removal tasks, and even slightly surpasses costly proprietary models.
Delong Liu, Haotian Hou, Zhaohui Hou, Shihao Han, Mingjie Zhan, Zhicheng Zhao 0001
AAAI4
2025 UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer
abstract
Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by aesthetic principles, we propose a non-invasive plug-in called Uniform Frame Organizer (UFO), which is compatible with any diffusion-based video generation model. The UFO comprises a series of adaptive adapters with adjustable intensities, which can significantly enhance the consistency between the foreground and background of videos and improve image quality without altering the original model parameters when integrated. The training for UFO is simple, efficient, requires minimal resources, and supports stylized training. Its modular design allows for the combination of multiple UFOs, enabling the customization of personalized video generation models. Furthermore, the UFO also supports direct transferability across different models of the same specification without the need for specific retraining. The experimental results indicate that UFO effectively enhances video generation quality and demonstrates its superiority in public video generation benchmarks.
Delong Liu, Zhaohui Hou, Mingjie Zhan, Shihao Han, Zhicheng Zhao 0001
AAAI4
2025 One-Minute Video Generation with Test-Time Training
abstract
Transformers today still struggle to generate one-minute videos because self-attention layers are inefficient for long context. Alternatives such as Mamba layers struggle to produce coherent scenes because their hidden states are small and less expressive. We experiment with Test-Time Training (TTT) layers, whose hidden states themselves can be neural networks, therefore larger and more expressive. Adding TTT layers into a pre-trained Transformer enables it to generate one-minute videos from text storyboards. We curate a dataset based on Tom and Jerry cartoons as a proof-of-concept benchmark. Compared to baselines such as Mamba 2, Gated DeltaNet, and sliding-window attention layers, TTT layers generate much more coherent videos that tell complete stories, leading by 34 Elo points in a human evaluation of 100 videos per method. Although promising, our results are still limited in physical realism, and the efficiency of our implementation can be further improved.Sample videos, code and annotations are available at: https://test-time-training.github.io/video-dit
Karan Dalal, Daniel Koceja, Yue Zhao 0006, Shihao Han, Ka Chun Cheung, Jan Kautz, Yejin Choi 0001, Yu Sun 0020, Xiaolong Wang 0004
CVPR5
2025 KAN-HyperpointNet for Point Cloud Sequence-Based 3D Human Action Recognition
abstract
Point cloud sequence-based 3D action recognition has achieved impressive performance and efficiency. However, existing point cloud sequence modeling methods cannot adequately balance the precision of limb micro-movements with the integrity of posture macro-structure, leading to the loss of crucial information cues in action inference. To overcome this limitation, we introduce D-Hyperpoint, a novel data type generated through a D-Hyperpoint Embedding module. D-Hyperpoint encapsulates both regional-momentary motion and global-static posture, effectively summarizing the unit human action at each moment. In addition, we present a D-Hyperpoint KANsMixer module, which is recursively applied to nested groupings of D-Hyperpoints to learn the action discrimination information and creatively integrates Kolmogorov-Arnold Networks (KAN) to enhance spatio-temporal interaction within D-Hyperpoints. Finally, we propose KAN-HyperpointNet, a spatio-temporal decoupled network architecture for 3D action recognition. Extensive experiments on two public datasets: MSR Action3D and NTU-RGB+D 60, demonstrate the state-of-the-art performance of our method.
Tianjin Yang, Shihao Han
ICASSP6
2025 Spatio-Temporal Point Convolutional Network With Meta-motion Level Refinement for Point Cloud-Based Human Action Recognition
abstract
Point cloud-based human action recognition leverages advanced spatio-temporal local encoding strategies to model the motion patterns, and has achieved remarkable performance. However, existing approaches struggle to precisely model the specific spatial scales and temporal spans of various body parts due to rigid spatio-temporal neighborhood limitations. Furthermore, due to the extraction of spatial configurations and temporal dynamics from isolated components, these methods cause distortions of spatio-temporal interactions. To solve the above problems, we propose a Meta-motion Level Refined Spatio-Temporal Point Convolution Network (MRST-PCN). Firstly, we progressively decouple the dynamic structures of different body parts, and present a logarithmic spatio-temporal point convolution strategy to capture meta-motion patterns at varying spatio-temporal granularities. Secondly, we devise the meta-motion differential strategy to derive short-term spatio-temporal displacements in frame neighborhoods, while designing a gated KANsformer to establish long-term spatio-temporal dependencies along the meta-motion flow. Finally, extensive experiments on three public datasets (MSR Action3D, UTD-MHAD, and NTU RGB+D 60) substantiate the superiority of MRST-PCN over state-of-the-art methods.
Shihao Han, Qing Meng
ICME4
2025 Geo-CF2Net: Geometry-Prior Cross-Frequency Interactive Fusion Network for 3D Human Action Recognition
abstract
Dynamic point cloud-based human action recognition has garnered increasing attention due to its inherent advantages in privacy preservation and structural completeness. Current methods typically rely on nested point spatio-temporal convolutions to understand motion semantics in a bottom-up manner, which is intractable for capturing high-fidelity human dynamics disentangled from spatio-temporal interference. Motivated by this, designing a practical spatio-temporal factorization backbone is essential. However, the repeated coarsening of aggregated features along the spatial dimension often leads to the degradation of intrinsic geometric texture relations within point cloud data. Moreover, discretizing continuous visual data into isolated temporal hyperpoints significantly diminishes temporal continuity, resulting in the fragmentation of human action. To circumvent above limitations, we propose a novel Geometry-Prior Cross-Frequency Interactive Fusion Network (Geo-CF2Net). Specifically, we investigate a Spatial-Geometry Pose Prior (SGPP) module, which compensates for pose information loss during spatial downsampling by explicitly modeling geometric constraints among neighboring points. In addition, we elaborate on a Temporal Motion Unit Interactive Coordination (TMIC) module to track the interactive composite semantics of low-frequency steady-state venations and high-frequency transient-state details within a high-dimensional pose evolution flow. Extensive experiments on three public benchmarks substantiate the superiority of Geo-CF2Net over state-of-the-art methods.
Qian Huang 0008, Xing Li 0005, Shihao Han, Yirui Wu, Xin Li 0090, Ziyang Yin
ACM Multimedia5
2025 FALCON: An ML Framework for Fully Automated Layout-Constrained Analog Circuit Design
abstract
Designing analog circuits from performance specifications is a complex, multi-stage process encompassing topology selection, parameter inference, and layout feasibility. We introduce FALCON, a unified machine learning framework that enables fully automated, specification-driven analog circuit synthesis through topology selection and layout-constrained optimization. Given a target performance, FALCON first selects an appropriate circuit topology using a performance-driven classifier guided by human design heuristics. Next, it employs a custom, edge-centric graph neural network trained to map circuit topology and parameters to performance, enabling gradient-based parameter inference through the learned forward model. This inference is guided by a differentiable layout cost, derived from analytical equations capturing parasitic and frequency-dependent effects, and constrained by design rules. We train and evaluate FALCON on a large-scale custom dataset of 1M analog mm-wave circuits, generated and simulated using Cadence Spectre across 20 expert-designed topologies. Through this evaluation, FALCON demonstrates >99\% accuracy in topology inference, <10\% relative error in performance prediction, and efficient layout-aware design that completes in under 1 second per instance. Together, these results position FALCON as a practical and extensible foundation model for end-to-end analog circuit design automation.
Asal Mehradfar, Xuzhe Zhao, Yilun Huang 0006, Emir Ceyani, Yankai Yang, Shihao Han, Hamidreza Aghasi, Amir Salman Avestimehr
NeurIPS6
2024 RNC: Efficient RRAM-aware NAS and Compilation for DNNs on Resource-Constrained Edge Devices
abstract
Computing-in-memory (CIM) is an emerging computing paradigm, offering noteworthy potential for accelerating neural networks with high parallelism, low latency, and energy efficiency compared to conventional von Neumann architectures. However, existing research has primarily focused on hardware architecture and network co-design for large-scale neural networks, without considering resource constraints. In this study, we aim to develop edge-friendly deep neural networks (DNNs) for accelerators based on resistive random-access memory (RRAM). To achieve this, we propose an edge compilation and resource-constrained RRAM-aware neural architecture search (NAS) framework to search for optimized neural networks meeting specific hardware constraints. Our compilation approach integrates layer partitioning, duplication, and network packing to maximize the utilization of computation units. The resulting network architecture can be optimized for either high accuracy or low latency using a one-shot neural network approach with Pareto optimality achieved through the Non-dominated Sorted Genetic Algorithm II (NSGA-II). The compilation of mobile-friendly networks, like Squeezenet and MobilenetV3 small can achieve over 80% of utilization and over 6x speedup compared to ISAAC-like framework with different crossbar resources. The resulting model from NAS optimized for speed achieved 5x-30x speedup. The code for this paper is available at https://github.com/ArChiiii/rram_nas_comp_pack.
Kam Chi Loong, Shihao Han, Sishuo Liu, Ning Lin
ICCD2
2024 CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts
abstract
Abstract Artificial intelligence (AI) has experienced substantial advancements recently, notably with the advent of large-scale language models (LLMs) employing mixture-of-experts (MoE) techniques, exhibiting human-like cognitive skills. As a promising hardware solution for edge MoE implementations, the computing-in-memory (CIM) architecture collocates memory and computing within a single device, significantly reducing the data movement and the associated energy consumption. However, due to diverse edge application scenarios and constraints, determining the optimal network structures for MoE, such as the expert’s location, quantity, and dimension on CIM systems remains elusive. To this end, we introduce a software-hardware co-designed neural architecture search (NAS) framework, C IM-based M oE N AS (CMN), focusing on identifying a high-performing MoE structure under specific hardware constraints. The results of the NYUD-v2 dataset segmentation on the RRAM (SRAM) CIM system reveal that CMN can discover optimized MoE configurations under energy, latency, and performance constraints, achieving 29.67 × ( 43.10 ×) energy savings, 175.44 ×( 109.89 ×) speedup, and 12.24 × smaller model size compared to the baseline MoE-enabled Visual Transformer, respectively. This co-design opens up an avenue toward high-performance MoE deployments in edge CIM systems.
Shihao Han, Sishuo Liu, Shucheng Du, Mingzi Li, Zijian Ye, Xiaoxin Xu, Dashan Shang
Sci. China Inf. Sci.1
2024 Erratum to: CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts
Shihao Han, Sishuo Liu, Shucheng Du, Mingzi Li, Zijian Ye, Xiaoxin Xu, Dashan Shang
Sci. China Inf. Sci.1
2021 Boosting Mobile CNN Inference through Semantic Memory
abstract
Human brains are known to be capable of speeding up visual recognition of repeatedly presented objects through faster memory encoding and accessing procedures on activated neurons. For the first time, we borrow and distill such a capability into a semantic memory design, namely SMTM, to improve on-device CNN inference. SMTM employs a hierarchical memory architecture to leverage the long-tail distribution of objects of interest, and further incorporates several novel techniques to put it into effects: (1) it encodes high-dimensional feature maps into low-dimensional, semantic vectors for low-cost yet accurate cache and lookup; (2) it uses a novel metric in determining the exit timing considering different layers' inherent characteristics; (3) it adaptively adjusts the cache size and semantic vectors to fit the scene dynamics. SMTM is prototyped on commodity CNN engine and runs on both mobile CPU and GPU. Extensive experiments on large-scale datasets and models show that SMTM can significantly speed up the model inference over standard approach (up to 2×) and prior cache designs (up to 1.5x), with acceptable accuracy loss.
Chen Zhang 0001, Shihao Han, Li Lyna Zhang, Baoqun Yin, Yunxin Liu 0001, Mengwei Xu 0001
ACM Multimedia3
2021 nn-Meter: towards accurate latency prediction of deep-learning model inference on diverse edge devices
abstract
With the recent trend of on-device deep learning, inference latency has become a crucial metric in running Deep Neural Network (DNN) models on various mobile and edge devices. To this end, latency prediction of DNN model inference is highly desirable for many tasks where measuring the latency on real devices is infeasible or too costly, such as searching for efficient DNN models with latency constraints from a huge model-design space. Yet it is very challenging and existing approaches fail to achieve a high accuracy of prediction, due to the varying model-inference latency caused by the runtime optimizations on diverse edge devices.
Li Lyna Zhang, Shihao Han, Jianyu Wei, Ningxin Zheng, Ting Cao 0003, Yuqing Yang 0001, Yunxin Liu 0001
MobiSys2