Shiyu Fan

dblp:302/0266 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot manipulation · 29% Generative modeling · 29% Face, body and person analysis · 25%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.012026
Generative Motion In-Betweening by Diffusion Over Continuous Implicit Representations · IEEE Trans. Vis. Comput. Graph. 2026
Machine learning › Generative modeling › diffusion model
latent diffusion model
1.012026
Generative Motion In-Betweening by Diffusion Over Continuous Implicit Representations · IEEE Trans. Vis. Comput. Graph. 2026
Computer animation and physical simulation › motion synthesis › motion interpolation
motion in-betweening
1.012026
Generative Motion In-Betweening by Diffusion Over Continuous Implicit Representations · IEEE Trans. Vis. Comput. Graph. 2026
Computer animation and physical simulation
motion synthesis
1.012026
Generative Motion In-Betweening by Diffusion Over Continuous Implicit Representations · IEEE Trans. Vis. Comput. Graph. 2026
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
0.912025
Waymo-3DSkelMo: A Multi-Agent 3D Skeletal Motion Dataset for Pedestrian Interaction Modeling in Autonomous Driving · ACM Multimedia 2025
Robotics › Robot manipulation
deformable object manipulation
0.912025
Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation · ICRA 2025
Robotics › Robot manipulation › deformable object manipulation
garment manipulation
0.912025
Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation · ICRA 2025
Computer vision › Face, body and person analysis › human pose estimation
human pose forecasting
0.912025
Waymo-3DSkelMo: A Multi-Agent 3D Skeletal Motion Dataset for Pedestrian Interaction Modeling in Autonomous Driving · ACM Multimedia 2025
Computer vision › 3D vision
implicit neural representation
0.312026
Generative Motion In-Betweening by Diffusion Over Continuous Implicit Representations · IEEE Trans. Vis. Comput. Graph. 2026
Robotics › Robot manipulation
grasping
0.312025
Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation · ICRA 2025

Methods — techniques the papers use, named apart from their topics

sampling optimization · 2.0latent diffusion · 2.0implicit neural representation · 2.0subtask decomposition · 0.9point cloud processing · 0.9multi-view RGB-D perception · 0.93d human body shape and motion priors · 0.9
YearPublicationVenuePosition
2026 VFE-CIM: An Algorithm-Hardware Co-Designed Computing-in-Memory Accelerator for Efficient Voxel Feature Encoding in Large-Scale Point Clouds
Yuang Ma, Shiyu Fan, Song Chen 0001, Yi Kang
ISCAS2
2026 I2Rec: Enabling Intra and Inter Batch Reuse in Recommendation Systems with PIM Architecture
abstract
The Deep Learning Recommendation Model (DLRM), one of the most popular recommendation system models, faces a performance bottleneck due to its memory-bound embedding layers. In recent years, processing-in-memory (PIM) has emerged as a solution to address the “memory wall” problem. Numerous PIM-based works have been published aiming to enhance DLRM performance by exploiting data locality. However, existing methods have yet to fully capitalize on locality. To better resolve the locality issue of the embedding layer and boost the performance of DLRM, we propose I2Rec, an architecture based on PIM that can further explore the locality in a DLRM system. I2Rec employs both intra-batch and inter-batch reuse strategies, releasing the potential of inter-batch reuse. As the embedding table size grows, I2Rec can uncover more reuse opportunities so the locality can be utilized more efficiently. Compared with spatial locality methods, I2Rec avoids a long preprocessing flow and achieves better locality exploration. Experimental results show that I2Rec achieves a 1.28× speedup and reduces memory accesses by 27% compared with intra-batch reuse alone under the same cache size and up to 1.41× speedup with a little extra overhead. Additionally, I2Rec outperforms state-of-the-art spatial locality algorithms, reducing memory traffic to 40% and achieving a 2.40× improvement in performance.
Shiyu Fan, Yuang Ma, Yi Kang
ACM Trans. Design Autom. Electr. Syst.1
2026 Generative Motion In-Betweening by Diffusion Over Continuous Implicit Representations
abstract
Recent advances in generative models have yielded impressive progress on motion in-betweening, allowing for more complex, varied, and realistic motion transitions. However, recent methods still exhibit noticeable limitations in preserving keyframe information and ensuring motion continuity. In this paper, we propose a novel pipeline and sampling optimization strategy for latent diffusion models (LDM) based on motion implicit neural representations (INR). By establishing a mapping between INR and sparse spatial or temporal information within latent diffusion, our model can sample the INR parameters from extremely sparse and ambiguous keyframe data and reconstruct plausible and smooth motions from the manifold. Our experiments demonstrate the superior performance of our model, which significantly improves motion generation quality in scenarios with few keyframes while ensuring both keyframe accuracy and diversity of in-between motions.
Shiyu Fan, Paul Henderson, Edmond S. L. Ho
IEEE Trans. Vis. Comput. Graph.1
2025 Leveraging LLMs for Automated Translation of Legacy Code: A Case Study on PL/SQL to Java Transformation
abstract
The VT legacy system, comprising approximately 2.5 million lines of PL/SQL code, lacks consistent documentation and automated tests, posing significant challenges for refactoring and modernisation. This study investigates the feasibility of leveraging large language models (LLMs) to assist in translating PL/SQL code into Java for the modernised "VTF3" system. By leveraging a dataset comprising 10 PL/SQL-to-Java code pairs and 15 Java classes, which collectively established a domain model for the translated files, multiple LLMs were evaluated. Furthermore, we propose a customized prompting strategy that integrates chain-of-guidance reasoning with n-shot prompting. Our findings indicate that this methodology effectively guides LLMs in generating syntactically accurate translations while also achieving functional correctness. However, the findings are limited by the small sample size of available code files and the restricted access to test cases used for validating the correctness of the generated code. Nevertheless, these findings lay the groundwork for scalable, automated solutions in modernising large legacy systems.
Lola Solovyeva, Eduardo Carneiro Oliveira, Shiyu Fan, Alper Tuncay, Shamil Gareev, Andrea Capiluppi
EASE3
2025 Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation
abstract
We present Flat'n'Fold, a novel large-scale dataset for garment manipulation that addresses critical gaps in existing datasets. Comprising 1,212 human and 887 robot demonstrations of flattening and folding 44 unique garments across 8 categories, Flat'n'Fold surpasses prior datasets in size, scope, and diversity. Our dataset uniquely captures the entire manipulation process from crumpled to folded states, providing synchronized multi-view RGB-D images, point clouds, and action data, including hand or gripper positions and rotations. We quantify the dataset's diversity and complexity compared to existing benchmarks and show that our dataset features natural and diverse manipulations of real-world demonstrations of human and robot demonstrations in terms of visual and action information. To showcase Flat'n'Fold's utility, we establish new benchmarks for grasping point prediction and subtask decomposition. Our evaluation of state-of-the-art models on these tasks reveals significant room for improvement. This underscores Flat'n'Fold's potential to drive advances in robotic perception and manipulation of deformable objects. Our dataset can be downloaded at https://cvas-ug.github.io/flat-n-fold
Lipeng Zhuang, Shiyu Fan, Yingdong Ru, Florent P. Audonnet, Paul Henderson, Gerardo Aragon-Camarasa
ICRA2
2025 Waymo-3DSkelMo: A Multi-Agent 3D Skeletal Motion Dataset for Pedestrian Interaction Modeling in Autonomous Driving
abstract
Large-scale high-quality 3D motion datasets with multi-person interactions are crucial for data-driven models in autonomous driving to achieve fine-grained pedestrian interaction understanding in dynamic urban environments. However, existing datasets mostly rely on estimating 3D poses from monocular RGB video frames, which suffer from occlusion and lack of temporal continuity, thus resulting in unrealistic and low-quality human motion. In this paper, we introduce Waymo-3DSkelMo, the first large-scale dataset providing high-quality, temporally coherent 3D skeletal motions with explicit interaction semantics, derived from the Waymo Perception dataset. Our key insight is to utilize 3D human body shape and motion priors to enhance the quality of the 3D pose sequences extracted from the raw LiDAR point clouds. The dataset covers over 14,000 seconds across more than 800 real driving scenarios, including rich interactions among an average of 27 agents per scene (with up to 250 agents in the largest scene). Furthermore, we establish 3D pose forecasting benchmarks under varying pedestrian densities, and the results demonstrate its value as a foundational resource for future research on fine-grained human behavior understanding in complex urban environments.
Guangxun Zhu, Shiyu Fan, Hang Dai, Edmond S. L. Ho
ACM Multimedia2
2025 Missing-modality enabled multi-modal fusion architecture for medical data
Muyu Wang, Shiyu Fan, Zhongrang Xie, Hui Chen 0010
J. Biomed. Informatics2
2025 APAV: An advanced pangenome analysis and visualization toolkit
abstract
Traditional pangenome analysis focuses on gene presence/absence variations (gene PAVs). However, the current methods for gene PAV analysis are insensitive to detect small but valuable mutations within gene regions, and they overlook variations in intergenic regions. Additionally, the visual inspection of PAVs is an important but time-consuming step for pangenome analysis and result interpretation. To address these issues, we present APAV, an advanced toolkit designed for comprehensive PAV analysis and visualization. It integrates gene element-level PAV analysis and provides PAV analysis for arbitrary given regions in a genome. The resulted PAV profile can be visualized and investigated interactively with reports in HTML format, enabling researchers to conveniently verify sequencing read depth, target region coverage, and intervals of absence for each PAV. Furthermore, APAV offers various subsequent analysis and visualization functions based on the PAV profile table, including basic statistics, sample clustering, genome size estimation, and phenotype association analysis. We demonstrated the capability of APAV with pangenome analysis of tumor genomes and rice genomes. Performing PAV analysis at the element level not only provides more accurate information about the variations but also uncovers a larger number of variations for the phenotype-genotype association studies. In the rice genome analysis, we identified over twenty thousand distributed genes and more than fifty thousand distributed genetic elements. In the tumor genome analysis, element-level analysis revealed approximately three times as many phenotype-related genes as gene-level analysis. This indicates that altering the PAV unit from genes to smaller segments or elements can lead to more biological insights.
Xiaorui Dong, Du Jiao, Hongzhang Xue, Shiyu Fan, Chaochun Wei
PLoS Comput. Biol.4
2024 Joint Local/Global Attention Cost Volume for Light Field Depth Estimation
abstract
Depth estimation of light field images played a significant role in various technology applications such as virtual reality, 3D modeling, and autonomous driving. However, existing deep learning methods tend to overlook the structural intricacies of the light field, leading to suboptimal performance in challenging areas like occlusion and textureless regions. Therefore, our paper proposes an attention cost volume network that combines local and global features to enhance performance in these challenging areas. We introduce a macro-pixel cost volume to effectively extract global context information, specifically targeting the challenge of objects in textureless regions. Then, our strategy combines attention cost volume from both local and global feature information to address the impact of occlusion. Finally, we introduce a new attention mechanism that generates attention weights to guide the cost volume. This mechanism is effective in eliminating redundancies and highlights crucial information, resulting in improved depth estimation quality. The experimental results on the HCI 4D light field dataset demonstrate that our proposed method exhibits smaller errors in occlusion and textureless regions compared to existing depth estimation methods.
Shiyu Fan, Huiping Deng, Sen Xiang
VCIP1
2022 Purification of tumor methylomes through residual decomposition
abstract
Due to the high heterogeneity of tumor tissue, methylation profiles of tumor samples obtained in clinical experiments are always mixture signals from different cellular components, including cancer, normal and stromal cells, etc. Among them, the admixture of normal cells is deemed as a major confounding factor for many downstream analyses. Decomposing mixture signals into profiles of their primitive constituents is vital for accurate differential calling and patient grouping. However, methods for purification of tumor methylomes are still lacking, even given a reliable estimate of tumor purity. In this work, we present ResDec, a residual-decomposition linear regression model for tumor methylome purification. We systematically evaluated the performance of our method compared with existing methods on both simulation data and TCGA methylation samples. ResDec achieves consistently better performance under different scenarios, including different numbers of matched normal samples, perturbations of input tumor purities and matched normal methylomes.
Nana Wei, Yijing Zhu, Yating Nie, Shiyu Fan, Yuanchen Sun, Xiaoqi Zheng
BIBM4