EDBT 2026 Demo / reviewers in the wild / expert
Mohan Xu
dblp:266/1188
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech SeparationabstractIn recent years, much speech separation research has focused primarily on improving model performance. However, for low-latency speech processing systems, high efficiency is equally important. Therefore, we propose a speech separation model with significantly reduced parameters and computational costs: Time-frequency Interleaved Gain Extraction and Reconstruction network (TIGER). TIGER leverages prior knowledge to divide frequency bands and compresses frequency information. We employ a multi-scale selective attention module to extract contextual features, while introducing a full-frequency-frame attention module to capture both temporal and frequency contextual information. Additionally, to more realistically evaluate the performance of speech separation models in complex acoustic environments, we introduce a dataset called EchoSet. This dataset includes noise and more realistic reverberation (e.g., considering object occlusions and material properties), with speech from two speakers overlapping at random proportions. Experimental results showed that models trained on EchoSet had better generalization ability than those trained on other datasets to the data collected in the physical world, which validated the practical value of the EchoSet. On EchoSet and real-world data, TIGER significantly reduces the number of parameters by 94.3% and the MACs by 95.3% while achieving performance surpassing state-of-the-art (SOTA) model TF-GridNet. Mohan Xu |
ICLR | 1 |
| 2025 | RaLMEN: A Robust UAV Recognition Framework for Low-Altitude Traffic SurveillanceabstractWith the rapid development of the low-altitude economy, Unmanned Aerial Vehicle (UAV) recognition has become increasingly important. However, vision-based methods progressively become ineffective as the UAV’s flight altitude increases. Therefore, radar information has emerged as the predominant approach for UAV recognition. However, existing low-altitude targets recognition methods either depend on long-sequence trajectory information or possess excessively high terminal computing power requirements, thereby rendering them difficult to deploy. In this study, a UAV recognition framework RaLMEN based on radar information is proposed, utilizing an ensemble learning architecture that incorporates BiLSTM and MLP. The framework integrates both temporal and static features extracted from aerial targets, aiming to provide a robust solution for low-altitude traffic surveillance. Experiments demonstrate that the proposed method requires only the temporal RCS information and position coordinates of a few trajectory points to accurately determine whether the target is a UAV. Furthermore, when evaluated with test data sampled from domains different from the training set, the framework demonstrates exceptional generalization performance, outperforming the baseline algorithms of all mainstream temporal neural networks. Mohan Xu, Gaoyuan Yang, Yuting Jia, Qiuyue Li, Qingyun Ye |
INDIN | 1 |
| 2024 | Domain Adaptation for Semantic Segmentation of Autonomous Driving with Contrastive LearningabstractSemantic segmentation is a critical component of autonomous driving perception systems and has gained increasing attention in recent advancements. Autonomous vehicles frequently encounter diverse environmental conditions, highlighting the significance of research into domain adaptation for semantic segmentation. We established a domain adversarial framework for enhancing the cross-domain perception of autonomous vehicles. However, previous works have shown that the general adversarial training-based methods can lead to indistinguishable features, resulting in a decline in the robustness of perception. In this regard, we adopted contrastive learning to guarantee the proximity of similar samples, and the principle of different classes of samples to a certain extent. This ensures the closeness of similar samples and upholds the feature distinction between different classes of samples to a significant degree. Thus, we proposed a novel domain adversarial training framework incorporating the contrastive learning method to enhance cross-domain feature recognition for autonomous driving systems. We empirically evaluate the proposed method against several recent baselines showing improved benchmark performances, confirming the effectiveness of the proposed method. Qiuyue Li, Mohan Xu, Bingtao Ren, Fan Zhou 0006 |
INDIN | 3 |
| 2023 | Hierarchical Transformer for Multi-Label Trailer Genre ClassificationabstractDetermining the genres of a trailer is a challenging multi-label classification task. Previous studies tend to classify by CNN or RNN. Recently, Transformer based on attention mechanism has achieved better results in many research fields than CNN and RNN. Inspired by these, we propose a Hierarchical Transformer (HT). HT can process both the frame sequence (HT-F) and audio (HT-A) of trailers. Besides, a feature compression module is inserted into HT-F, and audio spectrogram segment is processed by HT-A as a whole, which can effectively reduce the data processed by the second Transformer. In order to reduce the training cost and improve the performance, we load the pre-trained weights from other related fields into some parameters of HT, and utilize the limited resources to train the remaining parameters. Experiments show that our best model outperforms state-of-the-art methods on several comprehensive metrics. Zihui Cai, Xuemeng Wu, Mohan Xu, Xiaohui Cui |
ICASSP | 4 |
| 2023 | Application of deep reinforcement learning in attacking and protecting structural features-based malicious PDF detector
Xuemeng Wu, Mohan Xu, Xiaohui Cui |
Future Gener. Comput. Syst. | 4 |
| 2022 | Bandwidth-Efficient Multi-video Prefetching for Short Video StreamingabstractApplications that allow sharing of user-created short videos exploded in popularity in recent years. A typical short video application allows a user to swipe away the current video being watched and start watching the next video in a video queue. Such user interface causes significant bandwidth waste if users frequently swipe a video away before finishing watching. Solutions to reduce bandwidth waste without impairing the Quality of Experience (QoE) are needed. Solving the problem requires adaptively prefetching of short video chunks, which is challenging as the download strategy needs to match unknown user viewing behavior and network conditions. In our work, we first formulate the problem of adaptive multi-video prefetching in short video streaming. Then, to facilitate the integration and comparison of researchers' algorithms towards solving the problem, we design and implement a discrete-event simulator, which we release as open source. Finally, based on the organization of the Short Video Streaming Grand Challenge at ACM Multimedia 2022, we analyze and summarize the algorithms of the contestants, with the hope of promoting the research community towards addressing this problem. Xutong Zuo, Yishu Li, Mohan Xu, Wei Tsang Ooi, Jiangchuan Liu, Junchen Jiang, Xinggong Zhang, Kai Zheng 0003, Yong Cui 0001 |
ACM Multimedia | 3 |
| 2020 | Co-axial depth sensor with an extended depth range for AR/VR applicationsabstractDepth sensor is an essential element in virtual and augmented reality devices to digitalize users' environment in real time. The current popular technologies include the stereo, structured light, and Time-of-Flight (ToF). The stereo and structured light method require a baseline separation between multiple sensors for depth sensing, and both suffer from a limited measurement range. The ToF depth sensors have the largest depth range but the lowest depth map resolution. To overcome these problems, we propose a co-axial depth map sensor which is potentially more compact and cost-effective than conventional structured light depth cameras. Meanwhile, it can extend the depth range while maintaining a high depth map resolution. Also, it provides a high-resolution 2D image along with the 3D depth map. This depth sensor is constructed with a projection path and an imaging path. Those two paths are combined by a beamsplitter for a co-axial design. In the projection path, a cylindrical lens is inserted to add extra power in one direction which creates an astigmatic pattern. For depth measurement, the astigmatic pattern is projected onto the test scene, and then the depth information can be calculated from the contrast change of the reflected pattern image in two orthogonal directions. To extend the depth measurement range, we use an electronically focus tunable lens at the system stop and tune the power to implement an extended depth range without compromising depth resolution. In the depth measurement simulation, we project a resolution target onto a white screen which is moving along the optical axis and then tune the focus tunable lens power for three depth measurement subranges, namely, near, middle and far. In each sub-range, as the test screen moves away from the depth sensor, the horizontal contrast keeps increasing while the vertical contrast keeps decreasing in the reflected image. Therefore, the depth information can be obtained by computing the contrast ratio between features in orthogonal directions. The proposed depth map sensor could implement depth measurement for an extended depth range with a co-axial design. Mohan Xu, Hong Hua |
Virtual Real. Intell. Hardw. | 1 |