VLDB 2026 Research / reviewers in the wild / expert
Siyu Zou
dblp:324/6299
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Video understanding and tracking · 26% 3D vision · 26% Vision and language · 23% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model › multimodal large language model
unified image understanding and generation |
1.0 | 1 | 2026 | FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation · AAAI 2026 |
Visual content generation and editing › image editing
diffusion-based image editing |
0.8 | 1 | 2024 | Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks · AAAI 2024 |
Visual content generation and editing
image editing |
0.8 | 1 | 2024 | Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks · AAAI 2024 |
Visual content generation and editing › image editing
text-guided image editing |
0.8 | 1 | 2024 | Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks · AAAI 2024 |
Computer vision › 3D vision
3d human reconstruction |
0.6 | 1 | 2022 | HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic Consistency · ICRA 2022 |
Computer vision › Video understanding and tracking
action recognition |
0.6 | 1 | 2022 | PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action Recognition · ICRA 2022 |
Computer vision › 3D vision
human mesh recovery |
0.6 | 1 | 2022 | HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic Consistency · ICRA 2022 |
Computer vision › Video understanding and tracking › action recognition › multimodal action recognition
RGB-D action recognition |
0.6 | 1 | 2022 | PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action Recognition · ICRA 2022 |
Machine learning › Generative modeling
image generation |
0.3 | 1 | 2026 | FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation · AAAI 2026 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks · AAAI 2024 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.2 | 1 | 2024 | Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks · AAAI 2024 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.2 | 1 | 2022 | PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action Recognition · ICRA 2022 |
Machine learning › Generative modeling › trustworthy generative modeling
semantic consistency |
0.2 | 1 | 2022 | HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic Consistency · ICRA 2022 |
Methods — techniques the papers use, named apart from their topics
cross-modal attention · 1.5attention mask refinement · 1.5semantic-aware learning · 1.0transformer · 0.6parallel attention · 0.6cross-attention · 0.6adaptive weight fusion · 0.6CNN · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation
Wanggui He, Mushui Liu, Wenyi Xiao, Siyu Zou, Yanpeng Liu, Weilong Dai, Shuyi Ying, Ruikai Zhou, Yubo Tao, Hao Jiang 0062 |
AAAI | 5 |
| 2024 | Towards Efficient Diffusion-Based Image Editing with Instant Attention MasksabstractDiffusion-based Image Editing (DIE) is an emerging research hot-spot, which often applies a semantic mask to control the target area for diffusion-based editing. However, most existing solutions obtain these masks via manual operations or off-line processing, greatly reducing their efficiency. In this paper, we propose a novel and efficient image editing method for Text-to-Image (T2I) diffusion models, termed Instant Diffusion Editing (InstDiffEdit). In particular, InstDiffEdit aims to employ the cross-modal attention ability of existing diffusion models to achieve instant mask guidance during the diffusion steps. To reduce the noise of attention maps and realize the full automatics, we equip InstDiffEdit with a training-free refinement scheme to adaptively aggregate the attention distributions for the automatic yet accurate mask generation. Meanwhile, to supplement the existing evaluations of DIE, we propose a new benchmark called Editing-Mask to examine the mask accuracy and local editing ability of existing methods. To validate InstDiffEdit, we also conduct extensive experiments on ImageNet and Imagen, and compare it with a bunch of the SOTA methods. The experimental results show that InstDiffEdit not only outperforms the SOTA methods in both image quality and editing results, but also has a much faster inference speed, i.e., +5 to +6 times. Our code available at https://anonymous.4open.science/r/InstDiffEdit-C306 Siyu Zou, Jiji Tang, Yiyi Zhou, Chaoyi Zhao, Zhipeng Hu, Xiaoshuai Sun |
AAAI | 1 |
| 2022 | PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action RecognitionabstractDue to overly relying on appearance information or adopting direct static feature fusion, most of the existing action recognition methods based on multi-modality have poor robustness and insufficient consideration of modality differences. To address these problems, we propose a two-stream adaptive weight integration network with a three-dimensional parallel attention module, PA-AWCNN. Firstly, a three-dimensional Parallel Attention (PA) module is proposed to effectively extract features of spatial, temporal and channel dimensions and reduce the cross-dimensional interference, to achieve better robustness. Secondly, a Common Feature-driven (CFD) feature integration module is proposed to dynamically integrate appearance and depth features with adaptive weights, utilizing modality differences to redeem the lack of each feature, thereby balance the influence of both. The proposed PA-AW CNN uses the representative integrated feature generated by attention enhancement and feature integration for action recognition; it can not only get higher recognition accuracy but also improve the performance of distinguishing similar actions. Experiments illustrate that the proposed method achieves com-parable performances to state-of-the-art methods and obtains the accuracy of 92.76% and 95.65% on NTU RGB+D Dataset and SBU Kinect Interaction Dataset, respectively. The code is publicly available at: https://github.com/Luu-Yao/PA-AWCNN. Sheng Liu 0002, Chaonan Li, Siyu Zou, Shengyong Chen, Diyi Guan |
ICRA | 4 |
| 2022 | HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic ConsistencyabstractWe present a transformer-based network, Human Mesh Deformer (HMD-former), to tackle the problem of 3D human mesh reconstruction from a single RGB image. HMD-former applies a pre-trained CNN to extract image grid features and a transformer decoder to gradually warp the template 3D mesh to the deformed mesh. On each decoder layer, the fine-grained local information of grid features is well utilized using cross-attention by softly and content-dependently transforming the grid features to vertex embeddings. Auxiliary losses and proposed bi-directional mapping layers inherently ensure semantic consistency throughout the whole decoder, which free the network from learning unnecessary embedding transformation between layers. This further induces each layer of the decoder to focus on refining vertex embeddings and makes the whole network work in a progressively refining manner. Experiments on different public datasets Human3.6M and 3DPW show better reconstruction accuracy and faster inference speed than previous state-of-the-art methods, demonstrating the effectiveness and generalizability of HMD-former. Code is publicly available at https://github.com/siyuzou/HMD-former. Siyu Zou, Sheng Liu 0002, Chaonan Li, Shengyong Chen |
ICRA | 1 |