EDBT 2026 Demo / reviewers in the wild / expert
Zhuojie Wu
dblp:128/5757
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0005-7243-0901ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 68% Video understanding and tracking · 22% Autonomous driving · 5% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 50% Image and video coding · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › super-resolution
video super-resolution |
1.0 | 1 | 2026 | Compression-Oriented Video Super-Resolution · IEEE Trans. Image Process. 2026 |
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | 3DRealCar: An In-the-Wild RGB-D Car Dataset with 360-Degree Views · ICCV 2025 |
Computer vision › 3D vision
3d scene understanding |
0.9 | 1 | 2025 | 3DRealCar: An In-the-Wild RGB-D Car Dataset with 360-Degree Views · ICCV 2025 |
Computer vision › 3D vision › 3d reconstruction › multimodal 3d reconstruction
RGB-D reconstruction |
0.9 | 1 | 2025 | 3DRealCar: An In-the-Wild RGB-D Car Dataset with 360-Degree Views · ICCV 2025 |
Computer vision › 3D vision › 3d object recognition
multi-view recognition |
0.8 | 1 | 2024 | MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition Dataset · NeurIPS 2024 |
Computer vision › Video understanding and tracking
sign language recognition |
0.8 | 1 | 2024 | MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition Dataset · NeurIPS 2024 |
Computer vision › Video understanding and tracking
video propagation |
0.3 | 1 | 2026 | Compression-Oriented Video Super-Resolution · IEEE Trans. Image Process. 2026 |
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles |
0.3 | 1 | 2025 | 3DRealCar: An In-the-Wild RGB-D Car Dataset with 360-Degree Views · ICCV 2025 |
Computer vision › Face, body and person analysis
human pose |
0.2 | 1 | 2024 | MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition Dataset · NeurIPS 2024 |
Data stream processing
stream processing systems |
0.2 | 1 | 2013 | TimeStream: reliable stream computation in the cloud · EuroSys 2013 |
Distributed systems › stream processing
distributed stream processing |
0.2 | 1 | 2013 | TimeStream: reliable stream computation in the cloud · EuroSys 2013 |
Distributed systems › fault tolerance
failure recovery |
0.0 | 1 | 2013 | TimeStream: reliable stream computation in the cloud · EuroSys 2013 |
Distributed systems
fault tolerance |
0.0 | 1 | 2013 | TimeStream: reliable stream computation in the cloud · EuroSys 2013 |
Methods — techniques the papers use, named apart from their topics
motion vector refinement · 2.0metadata-driven alignment · 2.0compression-aware propagation · 2.0smartphone scanning · 0.9point cloud processing · 0.9multimodal fusion · 0.8cross-view benchmark · 0.8resilient substitution · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compression-Oriented Video Super-ResolutionabstractCurrent compressed video super-resolution methods have achieved promising performance, but they often assume that an input video is compressed under low-delay configurations. However, under random access configurations, those methods might struggle to leverage the metadata effectively due to the large variations of metadata in different compression configurations. In this work, we propose a Compression-Oriented Video Super-Resolution (COVSR) method that can address video super-resolution for both low-delay and random-access configurations. Specifically, we first introduce an efficient compression-aware propagation (ECAP) module that dynamically adjusts propagation routes in accordance with the compression configurations. Since existing methods require reconstructing frames in a frame-by-frame manner, it is difficult to achieve efficient parallelization. However, we find that by slightly relaxing sequential dependencies, our ECAP can significantly improve inference speed. Furthermore, existing methods typically perform alignment between adjacent frames or adjacent features. However, since ECAP may propagate features along non-adjacent reference routes, it introduces new challenges for accurate cross-frame feature alignment. In response, we propose a metadata-driven alignment (MDA) module that refines cross-frame motion vectors into dense, feature-level flow offsets, enabling precise alignment across temporally distant features. Extensive experimental results demonstrate that our COVSR not only achieves efficient and superior super-resolution performance but also is generalizable to various compression configurations. Our code will be available at https://covsr.github.io. Yanbin Liu 0003, Ming Lu 0002, Zhuojie Wu, Senmao Tian, Yandong Guo, Xin Yu 0002 |
IEEE Trans. Image Process. | 4 |
| 2025 | 3DRealCar: An In-the-Wild RGB-D Car Dataset with 360-Degree Viewsabstract3D cars are commonly used in self-driving systems, virtual/augmented reality, and games. However, existing 3D car datasets are either synthetic or low-quality, limiting their applications in practical scenarios and presenting a significant gap toward high-quality real-world 3D car datasets. In this paper, we propose the first large-scale 3D real car dataset, termed 3DRealCar, offering three distinctive features. (1) \textbf{High-Volume}: 2,500 cars are meticulously scanned by smartphones, obtaining car images and point clouds with real-world dimensions; (2) \textbf{High-Quality}: Each car is captured in an average of 200 dense, high-resolution 360-degree RGB-D views, enabling high-fidelity 3D reconstruction; (3) \textbf{High-Diversity}: The dataset contains various cars from over 100 brands, collected under three distinct lighting conditions, including reflective, standard, and dark. Additionally, we offer detailed car parsing maps for each instance to promote research in car parsing tasks. Moreover, we remove background point clouds and standardize the car orientation to a unified axis for the reconstruction only on cars and controllable rendering without background. We benchmark 3D reconstruction results with state-of-the-art methods across different lighting conditions in 3DRealCar. Extensive experiments demonstrate that the standard lighting condition part of 3DRealCar can be used to produce a large number of high-quality 3D cars, improving various 2D and 3D tasks related to cars. Notably, our dataset brings insight into the fact that recent 3D reconstruction methods face challenges in reconstructing high-quality 3D cars under reflective and dark lighting conditions. \textcolor{red}{\href{https://xiaobiaodu.github.io/3drealcar/}{Our dataset is here.}} Xiaobiao Du, Zhuojie Wu, Hongwei Sheng, Jiaying Ying, Ming Lu 0002, Tianqing Zhu, Kun Zhan, Xin Yu 0002 |
ICCV | 4 |
| 2024 | MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition DatasetabstractIsolated Sign Language Recognition (ISLR) focuses on identifying individual sign language glosses. Considering the diversity of sign languages across geographical regions, developing region-specific ISLR datasets is crucial for supporting communication and research. Auslan, as a sign language specific to Australia, still lacks a dedicated large-scale word-level dataset for the ISLR task. To fill this gap, we curate \underline{\textbf{the first}} large-scale Multi-view Multi-modal Word-Level Australian Sign Language recognition dataset, dubbed MM-WLAuslan. Compared to other publicly available datasets, MM-WLAuslan exhibits three significant advantages: (1) the largest amount of data, (2) the most extensive vocabulary, and (3) the most diverse of multi-modal camera views. Specifically, we record 282K+ sign videos covering 3,215 commonly used Auslan glosses presented by 73 signers in a studio environment.Moreover, our filming system includes two different types of cameras, i.e., three Kinect-V2 cameras and a RealSense camera. We position cameras hemispherically around the front half of the model and simultaneously record videos using all four cameras. Furthermore, we benchmark results with state-of-the-art methods for various multi-modal ISLR settings on MM-WLAuslan, including multi-view, cross-camera, and cross-view. Experiment results indicate that MM-WLAuslan is a challenging ISLR dataset, and we hope this dataset will contribute to the development of Auslan and the advancement of sign languages worldwide. All datasets and benchmarks are available at MM-WLAuslan. Heming Du, Hongwei Sheng, Hui Chen 0036, Huiqiang Chen, Zhuojie Wu, Xiaobiao Du, Jiaying Ying, Ruihan Lu, Qingzheng Xu, Xin Yu 0002 |
NeurIPS | 7 |
| 2024 | Exploring Generalizable Distillation for Efficient Medical Image SegmentationabstractEfficient medical image segmentation aims to provide accurate pixel-wise predictions with a lightweight implementation framework. However, existing lightweight networks generally overlook the generalizability of the cross-domain medical segmentation tasks. In this paper, we propose Generalizable Knowledge Distillation (GKD), a novel framework for enhancing the performance of lightweight networks on cross-domain medical segmentation by generalizable knowledge distillation from powerful teacher networks. Considering the domain gaps between different medical datasets, we propose the Model-Specific Alignment Networks (MSAN) to obtain the domain-invariant representations. Meanwhile, a customized Alignment Consistency Training (ACT) strategy is designed to promote the MSAN training. Based on the domain-invariant vectors in MSAN, we propose two generalizable distillation schemes, Dual Contrastive Graph Distillation (DCGD) and Domain-Invariant Cross Distillation (DICD). In DCGD, two implicit contrastive graphs are designed to model the intra-coupling and inter-coupling semantic correlations. Then, in DICD, the domain-invariant semantic vectors are reconstructed from two networks (i.e., teacher and student) with a crossover manner to achieve simultaneous generalization of lightweight networks, hierarchically. Moreover, a metric named Fréchet Semantic Distance (FSD) is tailored to verify the effectiveness of the regularized domain-invariant features. Extensive experiments conducted on the Liver, Retinal Vessel and Colonoscopy segmentation datasets demonstrate the superiority of our method, in terms of performance and generalization ability on lightweight networks. Xingqun Qi, Zhuojie Wu, Wenxuan Zou, Yifan Gao 0003, Muyi Sun, Shanghang Zhang, Caifeng Shan, Zhenan Sun |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | ShowFace: Coordinated Face Inpainting with Memory-Disentangled Refinement Networks
Zhuojie Wu, Xingqun Qi, Zijian Wang 0009, Kun Yuan 0003, Muyi Sun, Zhenan Sun |
BMVC | 1 |
| 2021 | PAENet: A Progressive Attention-Enhanced Network for 3D to 2D Retinal Vessel Segmentationabstract3D to 2D retinal vessel segmentation is a challenging problem in Optical Coherence Tomography Angiography (OCTA) images. Accurate retinal vessel segmentation is important for the diagnosis and prevention of ophthalmic diseases. However, making full use of the 3D data of OCTA volumes is a vital factor for obtaining satisfactory segmentation results. In this paper, we propose a Progressive Attention-Enhanced Network (PAENet) based on attention mechanisms to extract rich feature representation. Specifically, the framework consists of two main parts, the three-dimensional feature learning path and the two-dimensional segmentation path. In the three-dimensional feature learning path, we design a novel Adaptive Pooling Module (APM) and propose a new Quadruple Attention Module (QAM). The APM captures dependencies along the projection direction of volumes and learns a series of pooling coefficients for feature fusion, which efficiently reduces feature dimension. In addition, the QAM reweights the features by capturing four-group cross-dimension dependencies, which makes maximum use of 4D feature tensors. In the two-dimensional segmentation path, to acquire more detailed information, we propose a Feature Fusion Module (FFM) to inject 3D information into the 2D path. Meanwhile, we adopt the Polarized Self-Attention (PSA) block to model the semantic interdependencies in spatial and channel dimensions respectively. Experimentally, our extensive experiments on the OCTA-500 dataset show that our proposed algorithm achieves state-of-the-art performance compared with previous methods. Zhuojie Wu, Zijian Wang 0009, Wenxuan Zou, Fan Ji, Hao Dang, Muyi Sun |
BIBM | 1 |
| 2021 | CoCo DistillNet: a Cross-layer Correlation Distillation Network for Pathological Gastric Cancer SegmentationabstractIn recent years, deep convolutional neural networks have made significant advances in pathology image segmentation. However, pathology image segmentation encounters with a dilemma in which the higher-performance networks generally require more computational resources and storage. This phenomenon limits the employment of high-accuracy networks in real scenes due to the inherent high-resolution of pathological images. To tackle this problem, we propose CoCo DistillNet, a novel Cross-layer Correlation (CoCo) knowledge distillation network for pathological gastric cancer segmentation. Knowledge distillation, a general technique which aims at improving the performance of a compact network through knowledge transfer from a cumbersome network. Concretely, our CoCo DistillNet models the correlations of channel-mixed spatial similarity between different layers and then transfers this knowledge from a pre-trained cumbersome teacher network to a non-trained compact student network. In addition, we also utilize the adversarial learning strategy to further prompt the distilling procedure which is called Adversarial Distillation (AD). Furthermore, to stabilize our training procedure, we make the use of the unsupervised Paraphraser Module (PM) to boost the knowledge paraphrase in the teacher network. As a result, extensive experiments conducted on the Gastric Cancer Segmentation Dataset demonstrate the prominent ability of CoCo DistillNet which achieves state-of-the-art performance. Wenxuan Zou, Xingqun Qi, Zhuojie Wu, Zijian Wang 0009, Muyi Sun, Caifeng Shan |
BIBM | 3 |
| 2013 | TimeStream: reliable stream computation in the cloudabstractTimeStream is a distributed system designed specifically for low-latency continuous processing of big streaming data on a large cluster of commodity machines. The unique characteristics of this emerging application domain have led to a significantly different design from the popular MapReduce-style batch data processing. In particular, we advocate a powerful new abstraction called resilient substitution that caters to the specific needs in this new computation model to handle failure recovery and dynamic reconfiguration in response to load changes. Several real-world applications running on our prototype have been shown to scale robustly with low latency while at the same time maintaining the simple and concise declarative programming model. TimeStream handles an on-line advertising aggregation pipeline at a rate of 700,000 URLs per second with a 2-second delay, while performing sentiment analysis of Twitter data at a peak rate close to 10,000 tweets per second, with approximately 2-second delay. Zhengping Qian, Chunzhi Su, Zhuojie Wu, Hongyu Zhu 0003, Taizhi Zhang, Lidong Zhou, Zheng Zhang 0001 |
EuroSys | 4 |