Siyu Zou

dblp:324/6299 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Video understanding and tracking · 26% 3D vision · 26% Vision and language · 23%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model › multimodal large language model
unified image understanding and generation
1.012026
FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation · AAAI 2026
Visual content generation and editing › image editing
diffusion-based image editing
0.812024
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks · AAAI 2024
Visual content generation and editing
image editing
0.812024
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks · AAAI 2024
Visual content generation and editing › image editing
text-guided image editing
0.812024
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks · AAAI 2024
Computer vision › 3D vision
3d human reconstruction
0.612022
HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic Consistency · ICRA 2022
Computer vision › Video understanding and tracking
action recognition
0.612022
PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action Recognition · ICRA 2022
Computer vision › 3D vision
human mesh recovery
0.612022
HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic Consistency · ICRA 2022
Computer vision › Video understanding and tracking › action recognition › multimodal action recognition
RGB-D action recognition
0.612022
PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action Recognition · ICRA 2022
Machine learning › Generative modeling
image generation
0.312026
FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation · AAAI 2026
Machine learning › Generative modeling
diffusion model
0.212024
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks · AAAI 2024
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.212024
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks · AAAI 2024
Machine learning › Deep learning architectures and training
attention mechanism
0.212022
PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action Recognition · ICRA 2022
Machine learning › Generative modeling › trustworthy generative modeling
semantic consistency
0.212022
HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic Consistency · ICRA 2022

Methods — techniques the papers use, named apart from their topics

cross-modal attention · 1.5attention mask refinement · 1.5semantic-aware learning · 1.0transformer · 0.6parallel attention · 0.6cross-attention · 0.6adaptive weight fusion · 0.6CNN · 0.6
YearPublicationVenuePosition
2026 FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation
Wanggui He, Mushui Liu, Wenyi Xiao, Siyu Zou, Yanpeng Liu, Weilong Dai, Shuyi Ying, Ruikai Zhou, Yubo Tao, Hao Jiang 0062
AAAI5
2024 Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
abstract
Diffusion-based Image Editing (DIE) is an emerging research hot-spot, which often applies a semantic mask to control the target area for diffusion-based editing. However, most existing solutions obtain these masks via manual operations or off-line processing, greatly reducing their efficiency. In this paper, we propose a novel and efficient image editing method for Text-to-Image (T2I) diffusion models, termed Instant Diffusion Editing (InstDiffEdit). In particular, InstDiffEdit aims to employ the cross-modal attention ability of existing diffusion models to achieve instant mask guidance during the diffusion steps. To reduce the noise of attention maps and realize the full automatics, we equip InstDiffEdit with a training-free refinement scheme to adaptively aggregate the attention distributions for the automatic yet accurate mask generation. Meanwhile, to supplement the existing evaluations of DIE, we propose a new benchmark called Editing-Mask to examine the mask accuracy and local editing ability of existing methods. To validate InstDiffEdit, we also conduct extensive experiments on ImageNet and Imagen, and compare it with a bunch of the SOTA methods. The experimental results show that InstDiffEdit not only outperforms the SOTA methods in both image quality and editing results, but also has a much faster inference speed, i.e., +5 to +6 times. Our code available at https://anonymous.4open.science/r/InstDiffEdit-C306
Siyu Zou, Jiji Tang, Yiyi Zhou, Chaoyi Zhao, Zhipeng Hu, Xiaoshuai Sun
AAAI1
2022 PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action Recognition
abstract
Due to overly relying on appearance information or adopting direct static feature fusion, most of the existing action recognition methods based on multi-modality have poor robustness and insufficient consideration of modality differences. To address these problems, we propose a two-stream adaptive weight integration network with a three-dimensional parallel attention module, PA-AWCNN. Firstly, a three-dimensional Parallel Attention (PA) module is proposed to effectively extract features of spatial, temporal and channel dimensions and reduce the cross-dimensional interference, to achieve better robustness. Secondly, a Common Feature-driven (CFD) feature integration module is proposed to dynamically integrate appearance and depth features with adaptive weights, utilizing modality differences to redeem the lack of each feature, thereby balance the influence of both. The proposed PA-AW CNN uses the representative integrated feature generated by attention enhancement and feature integration for action recognition; it can not only get higher recognition accuracy but also improve the performance of distinguishing similar actions. Experiments illustrate that the proposed method achieves com-parable performances to state-of-the-art methods and obtains the accuracy of 92.76% and 95.65% on NTU RGB+D Dataset and SBU Kinect Interaction Dataset, respectively. The code is publicly available at: https://github.com/Luu-Yao/PA-AWCNN.
Sheng Liu 0002, Chaonan Li, Siyu Zou, Shengyong Chen, Diyi Guan
ICRA4
2022 HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic Consistency
abstract
We present a transformer-based network, Human Mesh Deformer (HMD-former), to tackle the problem of 3D human mesh reconstruction from a single RGB image. HMD-former applies a pre-trained CNN to extract image grid features and a transformer decoder to gradually warp the template 3D mesh to the deformed mesh. On each decoder layer, the fine-grained local information of grid features is well utilized using cross-attention by softly and content-dependently transforming the grid features to vertex embeddings. Auxiliary losses and proposed bi-directional mapping layers inherently ensure semantic consistency throughout the whole decoder, which free the network from learning unnecessary embedding transformation between layers. This further induces each layer of the decoder to focus on refining vertex embeddings and makes the whole network work in a progressively refining manner. Experiments on different public datasets Human3.6M and 3DPW show better reconstruction accuracy and faster inference speed than previous state-of-the-art methods, demonstrating the effectiveness and generalizability of HMD-former. Code is publicly available at https://github.com/siyuzou/HMD-former.
Siyu Zou, Sheng Liu 0002, Chaonan Li, Shengyong Chen
ICRA1