Zhiqi Ai

dblp:377/3743 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0005-1034-9972ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 StereoDETR: Stereo-Based Transformer for 3D Object Detection
abstract
Compared to monocular 3D object detection, stereo-based 3D methods offer significantly higher accuracy but still suffer from high computational overhead and latency. The state-of-the-art stereo 3D detection method achieves twice the accuracy of monocular approaches, yet its inference speed is only half as fast. In this paper, we propose StereoDETR, an efficient stereo 3D object detection framework based on DETR. StereoDETR consists of two branches: a monocular DETR branch and a stereo branch. The DETR branch is built upon 2D DETR with additional channels for predicting object scale, orientation, and sampling points. The stereo branch leverages low-cost multi-scale disparity features to predict object-level depth maps. These two branches are coupled solely through a differentiable depth sampling strategy. To handle occlusion, we introduce a constrained supervision strategy for sampling points without requiring extra annotations. Compared with the existing published monocular and binocular 3D detection methods, StereoDETR breaks the trade-off between speed and accuracy. Through a concise framework, it achieves binocular-level accuracy while maintaining monocular-level inference speed. The code is available at https://github.com/shiyi-mu/StereoDETR-OPEN.
Shiyi Mu, Zichong Gu, Zhiqi Ai, Shugong Xu
IEEE Trans. Circuits Syst. Video Technol.3
2025 VoxAging: Continuously Tracking Speaker Aging with a Large-Scale Longitudinal Dataset in English and Mandarin
Zhiqi Ai, Meixuan Bao, Xinnuo Li, Shugong Xu
INTERSPEECH1
2025 Towards Robust Speaker Recognition against Intrinsic Variation with Foundation Model Few-shot Tuning and Effective Speech Synthesis
Shuhang Wu, Xinnuo Li, Zhiqi Ai, Shugong Xu
INTERSPEECH4
2024 MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting
Zhiqi Ai, Shugong Xu
INTERSPEECH1
2024 StyleFusion TTS: Multimodal Style-Control and Enhanced Feature Fusion for Zero-Shot Text-to-Speech Synthesis
Xinnuo Li, Zhiqi Ai, Shugong Xu
PRCV (11)3
2024 Enhancing Open-Set Speaker Identification Through Rapid Tuning With Speaker Reciprocal Points and Negative Sample
abstract
This paper introduces a novel framework for open-set speaker identification in household environments, playing a crucial role in facilitating seamless human-computer interactions. Addressing the limitations of current speaker models and classification approaches, our work integrates an pretrained WavLM frontend with a few-shot rapid tuning neural network (NN) backend for enrollment, employing task-optimized Speaker Reciprocal Points Learning (SRPL) to enhance discrimination across multiple target speakers. Furthermore, we propose an enhanced version of SRPL (SRPL+), which incorporates negative sample learning with both speech-synthesized and real negative samples to significantly improve open-set SID accuracy. Our approach is thoroughly evaluated across various multi-language textdependent speaker recognition datasets, demonstrating its effectiveness in achieving high usability for complex household multi-speaker recognition scenarios. The proposed system enhanced open-set performance by up to 27% over the directly use of efficient WavLM base+ model. For detailed information on open-sourced implementation in our project website 1.
Zhiqi Ai, Xinnuo Li, Shugong Xu
SLT2