EDBT 2026 Demo / reviewers in the wild / expert
Junwei Zheng
dblp:296/6342
· DBLP profile ↗
22ranked-venue papers
3as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HybriDLA: Hybrid Generation for Document Layout AnalysisabstractConventional document layout analysis (DLA) traditionally depends on empirical priors or a fixed set of learnable queries executed in a single forward pass. While sufficient for early-generation documents with a small, predetermined number of regions, this paradigm struggles with contemporary documents, which exhibit diverse element counts and increasingly complex layouts. To address challenges posed by modern documents, we present HybriDLA, a novel generative framework that unifies diffusion and autoregressive decoding within a single layer. The diffusion component iteratively refines bounding-box hypotheses, whereas the autoregressive component injects semantic and contextual awareness, enabling precise region prediction even in highly varied layouts. To further enhance detection quality, we design a multi-scale feature-fusion encoder that captures both fine-grained and high-level visual cues. This architecture elevates performance to 83.5% mean Average Precision (mAP). Extensive experiments on the DocLayNet and M6Doc benchmarks demonstrate that HybriDLA sets a state-of-the-art performance, outperforming previous approaches. Yufan Chen 0001, Omar Moured, Ruiping Liu 0001, Junwei Zheng, Kunyu Peng, Jiaming Zhang 0001, Rainer Stiefelhagen |
AAAI | 4 |
| 2026 | Mitigating Label Noise using Prompt-Based Hyperbolic Meta-Learning in Open-Set Domain Generalization
Kunyu Peng, Di Wen 0006, M. Saquib Sarfraz, Yufan Chen 0001, Junwei Zheng, David Schneider 0006, Kailun Yang 0001, Alina Roitberg, Rainer Stiefelhagen |
Int. J. Comput. Vis. | 5 |
| 2025 | SAMBLE: Shape-Specific Point Cloud Sampling for an Optimal Trade-Off Between Local Detail and Global UniformityabstractDriven by the increasing demand for accurate and efficient representation of 3D data in various domains, point cloud sampling has emerged as a pivotal research topic in 3D computer vision. Recently, learning-to-sample methods have garnered growing interest from the community, particularly for their ability to be jointly trained with downstream tasks. However, previous learning-based sampling methods either lead to unrecognizable sampling patterns by generating a new point cloud or biased sampled results by focusing excessively on sharp edge details. Moreover, they all overlook the natural variations in point distribution across different shapes, applying a similar sampling strategy to all point clouds. In this paper, we propose a Sparse Attention Map and Bin-based Learning method (termed SAMBLE) to learn shape-specific sampling strategies for point cloud shapes. SAMBLE effectively achieves an improved balance between sampling edge points for local details and preserving uniformity in the global shape, resulting in superior performance across multiple common point cloud downstream tasks, even in scenarios with few-point sampling. Chengzhi Wu, Yuxin Wan, Julius Pfrommer, Zeyun Zhong, Junwei Zheng, Jürgen Beyerer |
CVPR | 6 |
| 2025 | Scene-agnostic Pose Regression for Visual LocalizationabstractAbsolute Pose Regression (APR) predicts 6D camera poses but lacks the adaptability to unknown environments without retraining, while Relative Pose Regression (RPR) generalizes better yet requires a large image retrieval database. Visual Odometry (VO) generalizes well in unseen environments but suffers from accumulated error in open trajectories. To address this dilemma, we introduce a new task, Scene-agnostic Pose Regression (SPR), which can achieve accurate pose regression in a flexible way while eliminating the need for retraining or databases. To benchmark SPR, we created a large-scale dataset, 360SPR, with over 200K photorealistic panoramas, 3.6M pinhole images and camera poses in 270 scenes at three different sensor heights. Furthermore, a SPR-Mamba model is initially proposed to address SPR in a dual-branch manner. Extensive experiments and studies demonstrate the effectiveness of our SPR paradigm, dataset, and model. In the unknown scenes of both 360SPR and 360Loc datasets, our method consistently outperforms APR, RPR and VO. The dataset and code are available at SPR. Junwei Zheng, Ruiping Liu 0001, Yufan Chen 0001, Zhenfang Chen, Kailun Yang 0001, Jiaming Zhang 0001, Rainer Stiefelhagen |
CVPR | 1 |
| 2025 | Graph-based Document Structure AnalysisabstractWhen reading a document, glancing at the spatial layout of a document is an initial step to understand it roughly. Traditional document layout analysis (DLA) methods, however, offer only a superficial parsing of documents, focusing on basic instance detection and often failing to capture the nuanced spatial and logical relationships between instances. These limitations hinder DLA-based models from achieving a gradually deeper comprehension akin to human reading. In this work, we propose a novel graph-based Document Structure Analysis (gDSA) task. This task requires that model not only detects document elements but also generates spatial and logical relations in form of a graph structure, allowing to understand documents in a holistic and intuitive manner. For this new task, we construct a relation graph-based document structure analysis dataset(GraphDoc) with 80K document images and 4.13M relation annotations, enabling training models to complete multiple tasks like reading order, hierarchical structures analysis, and complex inter-element relationship inference. Furthermore, a document relation graph generator (DRGG) is proposed to address the gDSA task, which achieves performance with 57.6% at $mAP_g$@$0.5$ for a strong benchmark baseline on this novel task and dataset. We hope this graphical representation of document structure can mark an innovative advancement in document structure analysis and understanding. The new dataset and code will be made publicly available. Yufan Chen 0001, Ruiping Liu 0001, Junwei Zheng, Di Wen 0006, Kunyu Peng, Jiaming Zhang 0001, Rainer Stiefelhagen |
ICLR | 3 |
| 2025 | Exploring Self-supervised Skeleton-based Action Recognition in Occluded EnvironmentsabstractTo integrate action recognition into autonomous robotic systems, it is essential to address challenges such as person occlusions—a common yet often overlooked scenario in existing self-supervised skeleton-based action recognition methods. In this work, we propose IosPSTL, a simple and effective self-supervised learning framework designed to handle occlusions. IosPSTL combines a cluster-agnostic KNN imputer with an Occluded Partial Spatio-Temporal Learning (OPSTL) strategy. First, we pre-train the model on occluded skeleton sequences. Then, we introduce a cluster-agnostic KNN imputer that performs semantic grouping using k-means clustering on sequence embeddings. It imputes missing skeleton data by applying K-Nearest Neighbors in the latent space, leveraging nearby sample representations to restore occluded joints. This imputation generates more complete skeleton sequences, which significantly benefits downstream self-supervised models. To further enhance learning, the OPSTL module incorporates Adaptive Spatial Masking (ASM) to make better use of intact, high-quality skeleton sequences during training. Our method achieves state-of-the-art performance on the occluded versions of the NTU-60 and NTU-120 datasets, demonstrating its robustness and effectiveness under challenging conditions. Code is available at https://github.com/cyfml/OPSTL. Kunyu Peng, Alina Roitberg, David Schneider 0006, Jiaming Zhang 0001, Junwei Zheng, Yufan Chen 0001, Ruiping Liu 0001, Kailun Yang 0001, Rainer Stiefelhagen |
IJCNN | 6 |
| 2025 | Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language ModelabstractPhysical environments and circumstances are fundamentally dynamic, yet current 3D datasets and evaluation benchmarks tend to concentrate on either dynamic scenarios or dynamic situations in isolation, resulting in incomplete comprehension. To overcome these constraints, we introduce Situat3DChange, an extensive dataset supporting three situation-aware change understanding tasks following the perception-action model: 121K question-answer pairs, 36K change descriptions for perception tasks, and 17K rearrangement instructions for the action task. To construct this large-scale dataset, Situat3DChange leverages 11K human observations of environmental changes to establish shared mental models and shared situational awareness for human-AI collaboration. These observations, enriched with egocentric and allocentric perspectives as well as categorical and coordinate spatial relations, are integrated using an LLM to support understanding of situated changes. To address the challenge of comparing pairs of point clouds from the same scene with minor changes, we propose SCReasoner, an efficient 3D MLLM approach that enables effective point cloud comparison with minimal parameter overhead and no additional tokens required for the language decoder. Comprehensive evaluation on Situat3DChange tasks highlights both the progress and limitations of MLLMs in dynamic scene and situation understanding. Additional experiments on data scaling and cross-domain transfer demonstrate the task-agnostic effectiveness of using Situat3DChange as a training dataset for MLLMs. The established dataset and source code are publicly available at: https://github.com/RuipingL/Situat3DChange. Ruiping Liu 0001, Junwei Zheng, Yufan Chen 0001, Kunyu Peng, Kailun Yang 0001, Jiaming Zhang 0001, Marc Pollefeys, Rainer Stiefelhagen |
NeurIPS | 2 |
| 2025 | HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person ScenariosabstractAction segmentation is a core challenge in high-level video understanding, aiming to partition untrimmed videos into segments and assign each a label from a predefined action set. Existing methods primarily address single-person activities with fixed action sequences, overlooking multi-person scenarios. In this work, we pioneer textual reference-guided human action segmentation in multi-person settings, where a textual description specifies the target person for segmentation. We introduce the first dataset for Referring Human Action Segmentation, i.e., RHAS133, built from 133 movies and annotated with 137 fine-grained actions with 33h video data, together with textual descriptions for this new task. Benchmarking existing action segmentation methods on RHAS133 using VLM-based feature extractors reveals limited performance and poor aggregation of visual cues for the target person. To address this, we propose a holistic-partial aware Fourier-conditioned diffusion framework, i.e., HopaDIFF, leveraging a novel cross-input gate attentional xLSTM to enhance holistic-partial long-range reasoning and a novel Fourier condition to introduce more fine-grained control to improve the action segmentation generation. HopaDIFF achieves state-of-the-art results on RHAS133 in diverse evaluation settings. The dataset and code are available at https://github.com/KPeng9510/HopaDIFF. Kunyu Peng, Junchao Huang, Xiangsheng Huang, Di Wen 0006, Junwei Zheng, Yufan Chen 0001, Kailun Yang 0001, Chongqing Hao, Rainer Stiefelhagen |
NeurIPS | 5 |
| 2025 | Exploring Video-Based Driver Activity Recognition under Noisy LabelsabstractAs an open research topic in the field of deep learning, learning with noisy labels has attracted much attention and grown rapidly over the past ten years. Learning with label noise is crucial for driver distraction behavior recognition, as real-world video data often contains mislabeled samples, impacting model reliability and performance. However, label noise learning is barely explored in the driver activity recognition field. In this paper, we propose the first label noise learning approach for the driver activity recognition task. Based on the cluster assumption, we initially enable the model to learn clustering-friendly low-dimensional representations from given videos and assign the resultant embeddings into clusters. We subsequently perform co-refinement within each cluster to smooth the classifier outputs. Furthermore, we propose a flexible sample selection strategy that combines two selection criteria without relying on any hyperparameters to filter clean samples from the training dataset. We also incorporate a self-adaptive parameter into the sample selection process to enforce balancing across classes. A comprehensive variety of experiments on the public Drive&Act dataset for all granularity levels demonstrates the superior performance of our method in comparison with other label-denoising methods derived from the image classification field. The source code is available at https://github.com/ilonafan/DAR-noisy-labels. Linjuan Fan, Di Wen 0006, Kunyu Peng, Kailun Yang 0001, Jiaming Zhang 0001, Ruiping Liu 0001, Yufan Chen 0001, Junwei Zheng, Rainer Stiefelhagen |
SMC | 8 |
| 2025 | Snap, Segment, Deploy: A Visual Data and Detection Pipeline for Wearable Industrial AssistantsabstractIndustrial assembly requires rapid adaptation to complex procedures under constrained computing, connectivity, and privacy conditions, rendering cloud-based solutions impractical. We present an on-device assistant for real-time, semi-hands-free guidance, integrating lightweight detection, speech recognition, and retrieval-augmented response generation. To enable scalable training without manual labeling, we construct the Gear8 dataset via an automated pipeline and introduce a two-stage refinement strategy to enhance robustness against domain shift. Experiments show improved generalization under diverse corruptions. User studies confirm notable gains in efficiency and error reduction, underscoring the system’s suitability for real-world industrial deployment. The Gear8 dataset, models, and code are publicly available at: https://github.com/Kratos-Wen/Gear8. Di Wen 0006, Junwei Zheng, Ruiping Liu 0001, Kunyu Peng, Rainer Stiefelhagen |
SMC | 2 |
| 2025 | @BENCH: Benchmarking Vision-Language Models for Human-centered Assistive TechnologyabstractAs Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However, benchmarking VLMs for ATs remains under-explored. To bridge this gap, we first create a novel AT benchmark (@ Bench). Guided by a pre-design user study with PVIs, our benchmark includes the five most crucial vision-language tasks: Panoptic Segmentation, Depth Estimation, Optical Character Recognition (OCR), Image Captioning, and Visual Question Answering (VQA). Besides, we propose a novel AT model (@MODEL) that addresses all tasks simultaneously and can be expanded to more assistive functions for helping PVIs. Our framework exhibits outstanding performance across tasks by integrating multi-modal information, and it offers PVIs a more comprehensive assistance. Extensive experiments prove the effectiveness and generalizability of our framework. Junwei Zheng, Ruiping Liu 0001, Jiaming Zhang 0001, Sven Matthiesen, Rainer Stiefelhagen |
WACV | 2 |
| 2024 | Navigating Open Set Scenarios for Skeleton-Based Action RecognitionabstractIn real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones. However, using pure skeleton data in such open-set conditions poses challenges due to the lack of visual background cues and the distinct sparse structure of body pose sequences. In this paper, we tackle the unexplored Open-Set Skeleton-based Action Recognition (OS-SAR) task and formalize the benchmark on three skeleton-based datasets. We assess the performance of seven established open-set approaches on our task and identify their limits and critical generalization issues when dealing with skeleton information.To address these challenges, we propose a distance-based cross-modality ensemble method that leverages the cross-modal alignment of skeleton joints, bones, and velocities to achieve superior open-set recognition performance. We refer to the key idea as CrossMax - an approach that utilizes a novel cross-modality mean max discrepancy suppression mechanism to align latent spaces during training and a cross-modality distance-based logits refinement method during testing. CrossMax outperforms existing approaches and consistently yields state-of-the-art results across all datasets and backbones. We will release the benchmark, code, and models to the community. Kunyu Peng, Junwei Zheng, Ruiping Liu 0001, David Schneider 0006, Jiaming Zhang 0001, Kailun Yang 0001, M. Saquib Sarfraz, Rainer Stiefelhagen, Alina Roitberg |
AAAI | 3 |
| 2024 | OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping
Jiale Wei, Junwei Zheng, Ruiping Liu 0001, Jie Hu 0039, Jiaming Zhang 0001, Rainer Stiefelhagen |
ACCV (10) | 2 |
| 2024 | RoDLA: Benchmarking the Robustness of Document Layout Analysis ModelsabstractBefore developing a Document Layout Analysis (DLA) model in real-world applications, conducting comprehensive robustness testing is essential. However, the robustness of DLA models remains underexplored in the literature. To address this, we are the first to introduce a robustness benchmark for DLA models, which includes 450K document images of three datasets. To cover realistic corruptions, we propose a perturbation taxonomy with 12 common document perturbations with 3 severity levels inspired by realworld document processing. Additionally, to better understand document perturbation impacts, we propose two metrics, Mean Perturbation Effect (mPE) for perturbation assessment and Mean Robustness Degradation (mRD) for robustness evaluation. Furthermore, we introduce a self-titled model, i.e., Robust Document Layout Analyzer (RoDLA), which improves attention mechanisms to boost extraction of robust features. Experiments on the proposed benchmarks (PubLayNet-P, DocLayNet-P, andM6Doc-P) demonstrate that RoDLA obtains state-of-the-art mRD scores of 115.7, 135.4, and 150.4, respectively. Compared to previous methods, RoDLA achieves notable improvements in mAP of +3.8%, +7.1% and +12.1%, respectively. Yufan Chen 0001, Jiaming Zhang 0001, Kunyu Peng, Junwei Zheng, Ruiping Liu 0001, Philip Torr 0001, Rainer Stiefelhagen |
CVPR | 4 |
| 2024 | Referring Atomic Video Action Recognition
Kunyu Peng, Jia Fu 0001, Kailun Yang 0001, Di Wen 0006, Yufan Chen 0001, Ruiping Liu 0001, Junwei Zheng, Jiaming Zhang 0001, M. Saquib Sarfraz, Rainer Stiefelhagen, Alina Roitberg |
ECCV (19) | 7 |
| 2024 | Open Panoramic Segmentation
Junwei Zheng, Ruiping Liu 0001, Yufan Chen 0001, Kunyu Peng, Chengzhi Wu, Kailun Yang 0001, Jiaming Zhang 0001, Rainer Stiefelhagen |
ECCV (39) | 1 |
| 2024 | Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-SupervisionabstractSelf-supervised representation learning for human action recognition has developed rapidly in recent years. Most of the existing works are based on skeleton data while using a multi-modality setup. These works overlooked the differences in performance among modalities, which led to the propagation of erroneous knowledge between modalities while only three fundamental modalities, i.e., joints, bones, and motions are used, hence no additional modalities are explored.In this work, we first propose an Implicit Knowledge Exchange Module (IKEM) which alleviates the propagation of erroneous knowledge between low-performance modalities. Then, we further propose three new modalities to enrich the complementary information between modalities. Finally, to maintain efficiency when introducing new modalities, we propose a novel teacher-student framework to distill the knowledge from the secondary modalities into the mandatory modalities considering the relationship constrained by anchors, positives, and negatives, named relational cross-modality knowledge distillation. The experimental results demonstrate the effectiveness of our approach, unlocking the efficient use of skeleton-based multi-modality data. Source code will be made publicly available at https://github.com/desehuileng0o0/IKEM. Yiping Wei, Kunyu Peng, Alina Roitberg, Jiaming Zhang 0001, Junwei Zheng, Ruiping Liu 0001, Yufan Chen 0001, Kailun Yang 0001, Rainer Stiefelhagen |
ICASSP | 5 |
| 2024 | Rethinking Attention Module Design for Point Cloud Analysis
Chengzhi Wu, Kaige Wang, Zeyun Zhong, Junwei Zheng, Julius Pfrommer, Jürgen Beyerer |
ICPR (26) | 5 |
| 2024 | MateRobot: Material Recognition in Wearable Robotics for People with Visual ImpairmentsabstractPeople with Visual Impairments (PVI) typically recognize objects through haptic perception. Knowing objects and materials before touching is desired by the target users but under-explored in the field of human-centered robotics. To fill this gap, in this work, a wearable vision-based robotic system, MATERobot, is established for PVI to recognize materials and object categories beforehand. To address the computational constraints of mobile platforms, we propose a lightweight yet accurate model MATEViT to perform pixel-wise semantic segmentation, simultaneously recognizing both objects and materials. Our methods achieve respective 40.2% and 51.1% of mIoU on COCOStuff-10K and DMS datasets, surpassing the previous method with +5.7% and +7.0% gains. Moreover, on the field test with participants, our wearable system reaches a score of 28 in the NASA-Task Load Index, indicating low cognitive demands and ease of use. Our MATERobot demonstrates the feasibility of recognizing material property through visual cues and offers a promising step towards improving the functionality of wearable robots for PVI. The source code has been made publicly available at MATERobot. Junwei Zheng, Jiaming Zhang 0001, Kailun Yang 0001, Kunyu Peng, Rainer Stiefelhagen |
ICRA | 1 |
| 2024 | Skeleton-Based Human Action Recognition with Noisy LabelsabstractUnderstanding human actions from body poses is critical for assistive robots sharing space with humans in order to make informed and safe decisions about the next interaction. However, precise temporal localization and annotation of activity sequences is time-consuming and the resulting labels are often noisy. If not effectively addressed, label noise negatively affects the model’s training, resulting in lower recognition quality. Despite its importance, addressing label noise for skeleton-based action recognition has been overlooked so far. In this study, we bridge this gap by implementing a framework that augments well-established skeleton-based human action recognition methods with label-denoising strategies from various research areas to serve as the initial benchmark. Observations reveal that these baselines yield only marginal performance when dealing with sparse skeleton data. Consequently, we introduce a novel methodology, NoiseEraSAR, which integrates global sample selection, co-teaching, and Cross-Modal Mixture-of-Experts (CM-MOE) strategies, aimed at mitigating the adverse impacts of label noise. Our proposed approach demonstrates better performance on the established benchmark, setting new state-of-the-art standards. The source code for this study will be made accessible at https://github.com/xuyizdby/NoiseEraSAR. Kunyu Peng, Di Wen 0006, Ruiping Liu 0001, Junwei Zheng, Yufan Chen 0001, Jiaming Zhang 0001, Alina Roitberg, Kailun Yang 0001, Rainer Stiefelhagen |
IROS | 5 |
| 2024 | Fourier Prompt Tuning for Modality-Incomplete Scene SegmentationabstractIntegrating information from multiple modalities enhances the robustness of scene perception systems in autonomous vehicles, providing a more comprehensive and reliable sensory framework. However, the modality incompleteness in multi-modal segmentation remains under-explored. In this work, we establish a task called Modality-Incomplete Scene Segmentation (MISS), which encompasses both system-level modality absence and sensor-level modality errors. To avoid the predominant modality reliance in multi-modal fusion, we introduce a Missing-aware Modal Switch (MMS) strategy to proactively manage missing modalities during training. Utilizing bit-level batch-wise sampling enhances the model’s performance in both complete and incomplete testing scenarios. Furthermore, we introduce the Fourier Prompt Tuning (FPT) method to incorporate representative spectral information into a limited number of learnable prompts that maintain robustness against all MISS scenarios. Akin to fine-tuning effects but with fewer tunable parameters (1.1%). Extensive experiments prove the efficacy of our proposed approach, showcasing an improvement of 5.84% mIoU over the prior state-of-the-art parameter-efficient methods in modality missing. The source code is publicly available at https://github.com/RuipingL/MISS. Ruiping Liu 0001, Jiaming Zhang 0001, Kunyu Peng, Yufan Chen 0001, Junwei Zheng, M. Saquib Sarfraz, Kailun Yang 0001, Rainer Stiefelhagen |
IV | 6 |
| 2023 | Attention-Based Point Cloud Edge SamplingabstractPoint cloud sampling is a less explored research topic for this data representation. The most commonly used sampling methods are still classical random sampling and farthest point sampling. With the development of neural networks, various methods have been proposed to sample point clouds in a task-based learning manner. However, these methods are mostly generative-based, rather than selecting points directly using mathematical statistics. Inspired by the Canny edge detection algorithm for images and with the help of the attention mechanism, this paper proposes a non-generative Attention-based Point cloud Edge Sampling method (APES), which captures salient points in the point cloud outline. Both qualitative and quantitative experimental results show the superior performance of our sampling method on common benchmark tasks. Chengzhi Wu, Junwei Zheng, Julius Pfrommer, Jürgen Beyerer |
CVPR | 2 |