EDBT 2026 Demo / reviewers in the wild / expert
Zixiao Yu
dblp:274/3789
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 33% Computer animation and physical simulation · 33% Multimedia systems and quality of experience · 33% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 100% | |
| Artificial intelligence
1 paper |
Graph learning · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning › graph neural network › attention-based graph neural network
graph attention network |
0.9 | 1 | 2025 | Gator: Accelerating Graph Attention Networks by Jointly Optimizing Attention and Graph Processing · ACM Trans. Archit. Code Optim. 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
graph neural network accelerator |
0.9 | 1 | 2025 | Gator: Accelerating Graph Attention Networks by Jointly Optimizing Attention and Graph Processing · ACM Trans. Archit. Code Optim. 2025 |
Visual content generation and editing › multimedia content creation
automated cinematography |
0.8 | 1 | 2024 | Automated Adaptive Cinematography for User Interaction in Open World · IEEE Trans. Multim. 2024 |
Computer animation and physical simulation › virtual cinematography
camera trajectory generation |
0.8 | 1 | 2024 | Automated Adaptive Cinematography for User Interaction in Open World · IEEE Trans. Multim. 2024 |
Multimedia systems and quality of experience
user interaction |
0.8 | 1 | 2024 | Automated Adaptive Cinematography for User Interaction in Open World · IEEE Trans. Multim. 2024 |
Methods — techniques the papers use, named apart from their topics
software-hardware co-design · 1.7parameter-adaptive feature selection · 1.7degree-weighted graph partitioning · 1.7relationship modeling · 0.8generative adversarial network · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gator: Accelerating Graph Attention Networks by Jointly Optimizing Attention and Graph ProcessingabstractGraph attention networks (GATs) have advanced performance in various application domains by introducing the attention mechanism into the graph neural networks (GNNs). The inefficiency of running GATs on CPUs or GPUs necessitates specialized hardware designs. Unfortunately, previous specialized architecture designs have focused on either the GNN architecture or the attention mechanism, resulting in limited performance and leaving ample room for improvement. This article presents Gator , a joint optimization approach with software–hardware co-designs for GAT inference. On the software level, Gator leverages degree-weighted graph partitioning and parameter-adaptive feature selection techniques to preprocess the input graph data, mining subgraph-level parallelism and mitigating the computation bottleneck of the dedicated dataflow. On the hardware level, Gator designs a unified processing engine to support various kernels by extracting a common computation pattern and a dimension-aware microarchitecture for efficient partial sum reduction. Extensive experiments show that our approach can achieve 11.5× more efficiency compared to NVIDIA RTX 4090 and provide a speedup of 3× to 9.4×, along with a 2.6× to 4.7× reduction in memory traffic, when compared to six state-of-the-art methods, with minimal accuracy loss. Xiaobo Lu, Jianbin Fang, Lin Peng 0001, Chun Huang 0006, Zixiao Yu |
ACM Trans. Archit. Code Optim. | 5 |
| 2024 | Automatic cinematography for body movement involved virtual communicationabstractAbstract The emergence of novel AI technologies and increasingly portable wearable devices have introduced a wider range of more liberated avenues for communication and interaction between human and virtual environments. In this context, the expression of distinct emotions and movements by users may convey a variety of meanings. Consequently, an emerging challenge is how to automatically enhance the visual representation of such interactions. Here, a novel Generative Adversarial Network (GAN) based model, AACOGAN, is introduced to tackle this challenge effectively. AACOGAN model establishes a relationship between player interactions, object locations, and camera movements, subsequently generating camera shots that augment player immersion. Experimental results demonstrate that AACOGAN enhances the correlation between player interactions and camera trajectories by an average of 73%, and improves multi‐focus scene quality up to 32.9%. Consequently, AACOGAN is established as an efficient and economical solution for generating camera shots appropriate for a wide range of interactive motions. Exemplary video footage can be found at https://youtu.be/Syrwbnpzgx8 . Zixiao Yu, Honghong Wang, Kim Un |
IET Commun. | 1 |
| 2024 | Automated Adaptive Cinematography for User Interaction in Open WorldabstractAdvancements in wearable technology and their capacity to interpret user movements, transforming them into interactive actions in virtual environments, have sparked an increased demand for user flexibility within these spaces. A direct outcome of this growing trend is the imperative need for automated cinematography in expansive, open-world scenarios. Nevertheless, the task of interpreting these interactive sequences through automated cinematography in unconstrained environments involves significant computational challenges. In response to this, we introduce the Automated Adaptive Cinematography for Open-world Generative Adversarial Network (AACOGAN) -an innovative solution that addresses these issues. Contrary to traditional models, which require comprehensive prior knowledge about scenes, characters, and objects, AACOGAN identifies and models the relationships among user interactions, object positions, and camera movements during the process of user engagement. This novel approach allows the model to function effectively even in open-world scenarios riddled with numerous uncertain factors. In the experimental phase, we developed and employed theMineStory Dataset, designed specifically for automatic cinematography in open-world scenarios. We devised and implemented novel metrics that are more congruent with the distinctive features of open-world scenarios. These innovative metrics provide a more nuanced understanding of the performance and effectiveness of our proposed method. Experimental findings substantiate that AACOGAN significantly enhances automatic cinematography performance within open-world contexts, including an average augmentation of 73% in the correlation between user interactions and camera trajectories, and an increase of up to 32.9% in the quality of multi-focus scenes. Therefore, AACOGAN emerges as an efficient, and innovative solution for creating appropriate camera shots in a myriad of interactive motions in open-world scenarios. An exemplary video footage can be found athttps://youtu.be/pbSHF-uxomw. Zixiao Yu, Xinyi Wu 0001, Haohong Wang, Aggelos K. Katsaggelos, Jian Ren 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Text Semantic Matching Research Based on Parallel Dropout
Zhuangzhuang Li, Zengzhen Shao, Jianxin Xiao, Zixiao Yu |
ICANN (5) | 4 |
| 2023 | A Novel Automatic Content Generation and Optimization FrameworkabstractWith the rapid growth of IoT multimedia devices, more and more content is being delivered in multimedia forms which generally requires more effort and resources from users to create and could be challenging to streamline. In this article, we propose Text2Animation (T2A), a framework that helps to generate complex multimedia content, and animation, from the simple textual script input. By leveraging recent advances in computational cinematography and video understanding, the purposed workflow further reduces associated knowledge requirements dramatically for cinematography. By jointly incorporating the fidelity and aesthetic models, T2A jointly considers the comprehensiveness of the visual presentation of the input script and the compliance of generated video with given cinematography specifications. The virtual camera placement in a 3-D environment is mapped into an optimization problem that can be resolved by using dynamic programming to achieve the lowest computational complexity. Experimental results show that T2A can reduce the manual animation production process by around 74%, and the new optimization framework can improve the perceptual quality of the output video by up to 35%. More video footage can be found athttps://www.youtube.com/watch?v=MMTJbmWL3gs. Zixiao Yu, Haohong Wang, Aggelos K. Katsaggelos, Jian Ren 0001 |
IEEE Internet Things J. | 1 |
| 2022 | RealPRNet: A Real-Time Phoneme-Recognized Network for "Believable" Speech AnimationabstractWith the technology development, more and more Internet of Things (IoT) devices with displays are making “face-to-face” interaction through visualization a reality. To protect the privacy of users, communications can be represented through avatars and use audio-driven real-time speech animation. However, if audio is the only available input, the quality of the outcome relies heavily on real-time phoneme recognition, such as recognition accuracy and latency. This article introduces a novel deep-learning-based real-time phoneme recognition network (RealPRNet) scheme to leverage spatial and temporal patterns in the input audio data. With featured long short-term memory stack block and long short-term features, RealPRNet can achieve super performance in phoneme recognition. Our comprehensive empirical results show that compared to the state-of-the-art algorithms, RealPRNet can achieve 20% phoneme error rate (PER) improvement and 4% block error distance (BDE) improvement in the best case. Zixiao Yu, Haohong Wang, Jian Ren 0001 |
IEEE Internet Things J. | 1 |