VLDB 2026 Research / reviewers in the wild / expert
Xiangwei Shi
dblp:232/5724
· DBLP profile ↗
5ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0003-0711-2147ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Face, body and person analysis · 67% 3D vision · 33% | |
| Computer graphics and multimedia
1 paper |
Computer animation and physical simulation · 77% Visual content generation and editing · 23% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer animation and physical simulation
motion synthesis |
0.9 | 1 | 2025 | SpeechCAT: Cross-Attentive Transformer for Audio to Motion Generation · HRI 2025 |
Computer vision › Face, body and person analysis
gaze estimation |
0.7 | 1 | 2023 | GazeNeRF: 3D-Aware Gaze Redirection with Neural Radiance Fields · CVPR 2023 |
Computer vision › Face, body and person analysis › gaze estimation
gaze redirection |
0.7 | 1 | 2023 | GazeNeRF: 3D-Aware Gaze Redirection with Neural Radiance Fields · CVPR 2023 |
Computer vision › 3D vision
neural radiance field |
0.7 | 1 | 2023 | GazeNeRF: 3D-Aware Gaze Redirection with Neural Radiance Fields · CVPR 2023 |
Visual content generation and editing
avatar generation |
0.3 | 1 | 2025 | SpeechCAT: Cross-Attentive Transformer for Audio to Motion Generation · HRI 2025 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9encoder-decoder · 0.9cross-attention · 0.9volume compositing · 0.7two-stream architecture · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SpeechCAT: Cross-Attentive Transformer for Audio to Motion GenerationabstractAudio-to-motion generation is an important task with applications in virtual avatar creation for XR systems and intelligent robot control in daily life scenarios. However, most existing motion generation methods rely on a single encoder-decoder architecture to model all body parts simultaneously, which limits their ability to capture the diverse and complex motions exhibited by humans. In this paper, we propose a novel method, SpeechCAT, that employs three separate encoder-decoder modules to individually model the motions of the face, body, and hands. To capture the relationships and synchronization among these body parts, we introduce a cross-attention mechanism to effectively learn their correlations. SpeechCAT ensures sufficient capacity to model the unique characteristics of each body part while preserving the coherence between them. Our experimental results demonstrate the superiority of SpeechCAT over baseline methods, highlighting its effectiveness in generating diverse, realistic, and synchronized motions with face, body, and hand parts. Sebastian Deaconu, Xiangwei Shi, Thomas Markhorst, Jouh Yeong Chew, Xucong Zhang |
HRI | 2 |
| 2023 | GazeNeRF: 3D-Aware Gaze Redirection with Neural Radiance FieldsabstractWe propose GazeNeRF, a 3D-aware method for the task of gaze redirection. Existing gaze redirection methods operate on 2D images and struggle to generate 3D consistent results. Instead, we build on the intuition that the face region and eyeballs are separate 3D structures that move in a coordinated yet independent fashion. Our method leverages recent advancements in conditional image-based neural radiance fields and proposes a two-stream architecture that predicts volumetric features for the face and eye regions separately. Rigidly transforming the eye features via a 3D rotation matrix provides fine-grained control over the desired gaze angle. The final, redirected image is then attained via differentiable volume compositing. Our experiments show that this architecture outperforms naively conditioned NeRF baselines as well as previous state-of-the-art 2D gaze redirection methods in terms of redirection accuracy and identity preservation. Code and models will be released for research purposes. Alessandro Ruzzi, Xiangwei Shi, Xi Wang 0021, Gengyan Li 0001, Shalini De Mello, Hyung Jin Chang, Xucong Zhang, Otmar Hilliges |
CVPR | 2 |
| 2020 | Zoom-CAM: Generating Fine-grained Pixel Annotations from Image LabelsabstractCurrent weakly supervised object localization and segmentation rely on class-discriminative visualization techniques to generate pseudo-labels for pixel-level training. Such visualization methods, including class activation mapping (CAM) and Grad-CAM, use only the deepest, lowest resolution convolutional layer, missing all information in intermediate layers. We propose Zoom-CAM: going beyond the last lowest resolution layer by integrating the importance maps over all activations in intermediate layers. Zoom-CAM captures fine-grained small-scale objects for various discriminative class instances, which are commonly missed by the baseline visualization methods. We focus on generating pixel-level pseudo-labels from class labels. The quality of our pseudo-labels evaluated on the ImageNet localization task exhibits more than 2.8% improvement on top-1 error. For weakly supervised semantic segmentation our generated pseudo-labels improve a state of the art model by 1.1%. Xiangwei Shi, Seyran Khademi, Yunqiang Li, Jan C. van Gemert |
ICPR | 1 |
| 2020 | WeightAlign: Normalizing Activations by Weight AlignmentabstractBatch normalization (BN) allows training very deep networks by normalizing activations by mini-batch sample statistics which renders BN unstable for small batch sizes. Current small-batch solutions such as Instance Norm, Layer Norm, and Group Norm use channel statistics which can be computed even for a single sample. Such methods are less stable than BN as they critically depend on the statistics of a single input sample. To address this problem, we propose a normalization of activation without sample statistics. We present WeightAlign: a method that normalizes the weights by the mean and scaled standard derivation computed within a filter, which normalizes activations without computing any sample statistics. Our proposed method is independent of batch size and stable over a wide range of batch sizes. Because weight statistics are orthogonal to sample statistics, we can directly combine WeightAlign with any method for activation normalization. We experimentally demonstrate these benefits for classification on CIFAR-10, CIFAR-100, ImageNet, for semantic segmentation on PASCAL VOC 2012 and for domain adaptation on Office-31. Xiangwei Shi, Yunqiang Li, Jan C. van Gemert |
ICPR | 1 |
| 2018 | Sight-Seeing in the Eyes of Deep Neural NetworksabstractWe address the interpretability of convolutional neural networks (CNNs) for predicting a geo-location from an image. In a pilot experiment we classify images of Pittsburgh vs Tokyo and visualize the learned CNN filters. We found that varying the CNN architecture leads to variating in the visualized filters. This calls for further investigation of the effective parameters on the interpretability of CNNs. Seyran Khademi, Xiangwei Shi, Tino Mager, Ronny Siebes, Carola Hein, Victor de Boer, Jan C. van Gemert |
eScience | 2 |