EDBT 2026 Demo / reviewers in the wild / expert
Zhuangzhuang Dai
dblp:267/4598
· DBLP profile ↗
8ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-6098-115XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ELLISON: An Advanced Multimodal Deep Fusion Framework for Attention Lapse Detection in Industrial Human-Robot CollaborationabstractAccurate evaluation of human attention in Human-Robot Collaboration (HRC) is essential to ensure intuitive and safe interactions. Although recent research has made progress in predicting human attention in social contexts, accurately estimating attention in industrial settings remains a challenge, particularly in industrial settings where it is crucial for error prevention, productivity optimization, and maintaining a secure work environment. In this paper, we present a multimodal deep fusion framework, ELLISON, designed to predict human attention lapses during HRC activities in manufacturing assembly tasks. First, we introduce a multimodal attention tracking pipeline featuring dual backbone Transformer Encoder models, which is trained to identify areas of interest related to human attention incorporating operator’s head pose and gaze features. Secondly, a deep fusion method is applied to consolidate the outputs of both modalities using a learned weighted strategy to provide a unified estimate of human attention throughout the tracking process. Experimental results demonstrate the effectiveness of our approach in tracking human attention in identifying instances where workers paid close attention or became distracted while working alongside robots, during collaborative tasks. The proposed method has potential applications in enhancing safety measures, optimizing workflows, and improving efficiency of HRC in industrial environments. Chen Li 0009, Zhuangzhuang Dai, Dimitrios Chrysostomou |
INDIN | 2 |
| 2025 | GazeTarget360: Towards Gaze Target Estimation in 360-Degree for Robot PerceptionabstractEnabling robots to understand human gaze target is a crucial step to allow capabilities in downstream tasks, for example, attention estimation and movement anticipation in real-world human-robot interactions. Prior works have addressed the in-frame target localization problem with data-driven approaches by carefully removing out-of-frame samples. Vision-based gaze estimation methods, such as OpenFace, do not effectively absorb background information in images and cannot predict gaze target in situations where subjects look away from the camera. In this work, we propose a system to address the problem of 360-degree gaze target estimation from an image in generalized visual scenes. The system, named GazeTarget360, integrates conditional inference engines of an eye-contact detector, a pre-trained vision encoder, and a multi-scale-fusion decoder. Cross validation results show that GazeTarget360 can produce accurate and reliable gaze target predictions in unseen scenarios. This makes a first-of-its-kind system to predict gaze targets from realistic camera footage which is highly efficient and deployable. Our source code is made publicly available at: https://github.com/zdai257/DisengageNet. Zhuangzhuang Dai, Vincent Gbouna Zakka, Luis Manso, Chen Li 0009 |
IROS | 1 |
| 2024 | SoundCount: Sound Counting from Raw Audio with Dyadic Decomposition Neural NetworkabstractIn this paper, we study an underexplored, yet important and challenging problem: counting the number of distinct sounds in raw audio characterized by a high degree of polyphonicity. We do so by systematically proposing a novel end-to-end trainable neural network~(which we call DyDecNet, consisting of a dyadic decomposition front-end and backbone network), and quantifying the difficulty level of counting depending on sound polyphonicity. The dyadic decomposition front-end progressively decomposes the raw waveform dyadically along the frequency axis to obtain time-frequency representation in multi-stage, coarse-to-fine manner. Each intermediate waveform convolved by a parent filter is further processed by a pair of child filters that evenly split the parent filter's carried frequency response, with the higher-half child filter encoding the detail and lower-half child filter encoding the approximation. We further introduce an energy gain normalization to normalize sound loudness variance and spectrum overlap, and apply it to each intermediate parent waveform before feeding it to the two child filters. To better quantify sound counting difficulty level, we further design three polyphony-aware metrics: polyphony ratio, max polyphony and mean polyphony. We test DyDecNet on various datasets to show its superiority, and we further show dyadic decomposition network can be used as a general front-end to tackle other acoustic tasks. Zhuangzhuang Dai, Agathoniki Trigoni, Long Chen 0005, Andrew Markham |
AAAI | 2 |
| 2024 | EgoCap and EgoFormer: First-person image captioning with context fusion
Zhuangzhuang Dai, Andrew Markham, Agathoniki Trigoni, M. Arif Imtiazur Rahman, Lahiru N. S. Wijayasingha, John A. Stankovic, Chen Li 0009 |
Pattern Recognit. Lett. | 1 |
| 2024 | Illumination-Aware Hallucination-Based Domain Adaptation for Thermal Pedestrian DetectionabstractThermal imagery is emerging as a viable candidate for 24-7, all-weather pedestrian detection owning to thermal sensors’ robust performance for pedestrian detection under different weather and illumination conditions. Despite the promising results obtained from combining visible (RGB) and thermal cameras in multi-spectral fusion techniques, the complex synchronization requirements, including alignment and calibration of sensors, impede their deployment in real-world scenarios. In this paper, we introduce a novel approach for domain adaptation to enhance the performance of pedestrian detection based solely on thermal images. Our proposed approach involves several stages. Firstly, we use both thermal and visible images as input during the training phase. Secondly, we leverage a thermal-to-visible hallucination network to generate feature maps that are similar to those generated by the visible branch. Finally, we design a transformer-based multi-modal fusion module to integrate the hallucinated visible and thermal information more effectively. The thermal-to-visible hallucination network acts as domain adaptation, allowing us to obtain pseudo-visual and thermal features using solely thermal input. Based on the experimental results, it is observed the mean average precision (mAP) increases by 4.72% and the miss rate decreases by 7.56% on the KAIST dataset when compared to the baseline model. Qian Xie 0001, Ta Ying Cheng, Zhuangzhuang Dai, Vu H. Tran, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | OdomBeyondVision: An Indoor Multi-modal Multi-platform Odometry Dataset Beyond the Visible SpectrumabstractThis paper presents a multimodal indoor odometry dataset, OdomBeyondVision, featuring multiple sensors across the different spectrum and collected with different mobile platforms. Not only does OdomBeyondVision contain the traditional navigation sensors, sensors such as IMUs, mechanical LiDAR, RGBD camera, it also includes several emerging sensors such as the single-chip mmWave radar, LWIR thermal camera and solid-state LiDAR. With the above sensors on UAV, UGV and handheld platforms, we respectively recorded the multimodal odometry data and their movement trajectories in various indoor scenes and different illumination conditions. We release the exemplar radar, radar-inertial and thermal-inertial odometry implementations to demonstrate their results for future works to compare against and improve upon. The full dataset including toolkit and documentation is publicly available at: https://github.com/MAPS-Lab/OdomBeyondVision. Peize Li, Kaiwen Cai, Muhamad Risqi Utama Saputra, Zhuangzhuang Dai, Xiaoxuan Lu 0001 |
IROS | 4 |
| 2022 | DeepCIR: Insights into CIR-based Data-driven UWB Error MitigationabstractUltra-Wide-Band (UWB) ranging sensors have been widely adopted for robotic navigation thanks to their extremely high bandwidth and hence high resolution. However, off-the-shelf devices may output ranges with significant errors in cluttered, severe non-line-of-sight (NLOS) environments. Recently, neural networks have been actively studied to improve the ranging accuracy of UWB sensors using the channel-impulse-response (CIR) as input. However, previous works have not systematically evaluated the efficacy of various packet types and their possible combinations in a two-way-ranging transaction, including poll, response and final packets. In this paper, we firstly investigate the utility of different packet types and their combinations when used as input for a neural network. Secondly, we propose two novel data-driven approaches, namely FMCIR and WMCIR, that leverage two-sided CIRs for efficient UWB error mitigation. Our approaches outperform state-of-the-art by a significant margin, further reducing range errors up to 45%. Finally, we create and release a dataset of transaction-level synchronized CIRs (each sample consists of the CIR of the poll, response and final packets), which will enable further studies in this area. Zhuangzhuang Dai, Agathoniki Trigoni, Andrew Markham |
IROS | 2 |
| 2020 | Indoor positioning system in visually-degraded environments with millimetre-wave radar and inertial sensors: demo abstractabstractPositional estimation is of great importance in the public safety sector. Emergency responders such as fire fighters, medical rescue teams, and the police will all benefit from a resilient positioning system to deliver safe and effective emergency services. Unfortunately, satellite navigation (e.g., GPS) offers limited coverage in indoor environments. It is also not possible to rely on infrastructure based solutions. To this end, wearable sensor-aided navigation techniques, such as those based on camera and Inertial Measurement Units (IMU), have recently emerged recently as an accurate, infrastructure-free solution. Together with an increase in the computational capabilities of mobile devices, motion estimation can be performed in real-time. In this demonstration, we present a real-time indoor positioning system which fuses millimetre-wave (mmWave) radar and IMU data via deep sensor fusion. We employ mmWave radar rather than an RGB camera as it provides better robustness to visual degradation (e.g., smoke, darkness, etc.) while at the same time requiring lower computational resources to enable runtime computation. We implemented the sensor system on a handheld device and a mobile computer running at 10 FPS to track a user inside an apartment. Good accuracy and resilience were exhibited even in poorly illuminated scenes. Zhuangzhuang Dai, Muhamad Risqi Utama Saputra, Xiaoxuan Lu 0001, Agathoniki Trigoni, Andrew Markham |
SenSys | 1 |