EDBT 2026 Demo / reviewers in the wild / expert
Jinzheng Zhao
dblp:253/1912
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2025
0009-0004-3910-850XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Textless Streaming Speech-to-Speech Translation using Semantic Speech TokensabstractCascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumulate. In this paper, we propose a transducer-based speech translation model that outputs discrete speech tokens in a low-latency streaming fashion. This approach eliminates the need for generating text output first, followed by machine translation (MT) and text-to-speech (TTS) systems. The produced speech tokens can be directly used to generate a speech signal with low latency by utilizing an acoustic language model (LM) to obtain acoustic tokens and an audio codec model to retrieve the waveform. Experimental results show that the proposed method outperforms other existing approaches and achieves state-of-the-art results for streaming translation in terms of BLEU, average latency, and BLASER 2.0 scores for multiple language pairs using the CVSS-C dataset as a benchmark. Jinzheng Zhao, Niko Moritz, Egor Lakomkin, Ruiming Xie, Zhiping Xiu, Katerina Zmolíková, Yashesh Gaur, Christian Fügen |
ICASSP | 1 |
| 2024 | Fusion of Audio and Visual Embeddings for Sound Event Localization and DetectionabstractSound event localization and detection (SELD) combines two subtasks: sound event detection (SED) and direction of arrival (DOA) estimation. SELD is usually tackled as an audio-only problem, but visual information has been recently included. Few audio-visual (AV)-SELD works have been published and most employ vision via face/object bounding boxes, or human pose keypoints. In contrast, we explore the integration of audio and visual feature embeddings extracted with pre-trained deep networks. For the visual modality, we tested ResNet50 and Inflated 3D ConvNet (I3D). Our comparison of AV fusion methods includes the AV-Conformer and Cross-Modal Attentive Fusion (CMAF) model. Our best models outperform the DCASE 2023 Task3 audio-only and AV baselines by a wide margin on the development set of the STARSS23 dataset, making them competitive amongst state-of-the-art results of the AV challenge, without model ensembling, heavy data augmentation, or prediction post-processing. Such techniques and further pre-training could be applied as next steps to improve performance. Davide Berghi, Peipei Wu, Jinzheng Zhao, Wenwu Wang 0001, Philip J. B. Jackson |
ICASSP | 3 |
| 2024 | ForecasterFlexOBM: A Multi-View Audio-Visual Dataset for Flexible Object-Based Media ProductionabstractLeveraging machine learning techniques, in the context of object-based media production, could enable provision of personalized media experiences to diverse audiences. To fine-tune and evaluate techniques for personalization applications, as well as more broadly, datasets which bridge the gap between research and production are needed. We introduce and release such a dataset, themed around a UK weather forecast and shot against a blue-screen background, of three professional actors/presenters – one male and one female (English) and one female (British Sign Language). Scenes include both production and research-oriented examples, with a range of dialogues and actions. Capture techniques consisted of a synchronized 4K resolution 16-camera array, production-typical microphones plus professional audio mix, a 16-channel microphone array with collocated Grasshopper3 camera, and a photogrammetry array. We demonstrate applications relevant to virtual production and creation of personalized media including neural radiance fields, shadow casting, action/event detection, speaker source tracking and video captioning. Davide Berghi, Craig Cieciura, Farshad Einabadi, Maxine Glancy, Oliver C. Camilleri, Philip Foster, Asmar Nadeem, Faegheh Sardari, Jinzheng Zhao, Marco Volino, Armin Mustafa, Philip J. B. Jackson, Adrian Hilton 0001 |
ICME | 9 |
| 2022 | PSSAT: A Perturbed Semantic Structure Awareness Transferring Method for Perturbation-Robust Slot FillingabstractMost existing slot filling models tend to memorize inherent patterns of entities and corresponding contexts from training data. However, these models can lead to system failure or undesirable outputs when being exposed to spoken language perturbation or variation in practice. We propose a perturbed semantic structure awareness transferring method for training perturbation-robust slot filling models. Specifically, we introduce two MLM-based training strategies to respectively learn contextual semantic structure and word distribution from unsupervised language perturbation corpus. Then, we transfer semantic knowledge learned from upstream training procedure into the original samples and filter generated data by consistency processing. These procedures aims to enhance the robustness of slot filling models. Experimental results show that our method consistently outperforms the previous basic methods and gains strong generalization while preventing the model from memorizing inherent patterns of entities and contexts. Guanting Dong 0001, Daichi Guo, Liwen Wang 0007, Xuefeng Li 0002, Zechen Wang, Keqing He 0001, Jinzheng Zhao, Yi Huang 0017, Junlan Feng, Weiran Xu |
COLING | 8 |
| 2022 | A Robust Contrastive Alignment Method for Multi-Domain Text ClassificationabstractMulti-domain text classification can automatically classify texts in various scenarios. Due to the diversity of human languages, texts with the same label in different domains may differ greatly, which brings challenges to the multi-domain text classification. Current advanced methods use the private-shared paradigm, capturing domain-shared features by a shared encoder, and training a private encoder for each domain to extract domain-specific features. However, in realistic scenarios, these methods suffer from inefficiency as new domains are constantly emerging. In this paper, we propose a robust contrastive alignment method to align text classification features of various domains in the same feature space by supervised contrastive learning. By this means, we only need two universal feature extractors to achieve multi-domain text classification. Extensive experimental results show that our method performs on par with or sometimes better than the state-of-the-art method, which uses the complex multi-classifier in a private-shared framework. Xuefeng Li 0002, Liwen Wang 0007, Guanting Dong 0001, Jinzheng Zhao, Jiachi Liu, Weiran Xu, Chunyun Zhang |
ICASSP | 5 |
| 2022 | Partial Arithmetic Consensus based Distributed Intensity Particle Flow SMC-PHD Filter for Multi-Target TrackingabstractIntensity Particle Flow (IPF) SMC-PHD has been proposed recently for multi-target tracking. In this paper, we extend IPF-SMC-PHD filter to distributed setting, and develop a novel consensus method for fusing the estimates from individual sensors, based on Arithmetic Average (AA) fusion. Different from conventional AA method which may be degraded when unreliable estimates are presented, we develop a novel arithmetic consensus method to fuse estimates from each individual IPF-SMC-PHD filter with partial consensus. The proposed method contains a scheme for evaluating the reliability of the sensor nodes and preventing unreliable sensor information to be used in fusion and communication in sensor network, which help improve fusion accuracy and reduce sensor communication costs. Numerical simulations are performed to demonstrate the advantages of the proposed algorithm over the uncooperative IPF-SMC-PHD and distributed particle-PHD with AA fusion. Peipei Wu, Jinzheng Zhao, Shidrokh Goudarzi, Wenwu Wang 0001 |
ICASSP | 2 |
| 2022 | Audio-Visual Tracking of Multiple Speakers Via a PMBM FilterabstractAudio-visual tracking of multiple speakers requires to estimate the state (e.g. velocity and location) of each speaker by leveraging the information of both audio and visual modalities. Estimating the number of speakers and their states jointly remains a challenging problem. We propose an Audio-Visual Possion Multi-Bernoulli Mixture Filter (AV-PMBM) that can not only predict the number of speakers but also give accurate estimation of their states. We also propose a novel sound source localization technique based on DOA information and a deep learning based object detector to provide reliable audio measurements for the AV tracker. To our knowledge, this represents the first attempt using PMBM for multi-speaker tracking with audio visual modalities. Experiments on the AV16.3 dataset demonstrate that AV-PMBM achieves state-of-the-art performance in optimal sub-pattern assignment (OSPA). Jinzheng Zhao, Peipei Wu, Xubo Liu 0001, Yong Xu 0004, Lyudmila Mihaylova, Simon J. Godsill, Wenwu Wang 0001 |
ICASSP | 1 |
| 2022 | Separate What You Describe: Language-Queried Audio Source SeparationabstractIn this paper, we introduce the task of language-queried audio source separation (LASS), which aims to separate a target source from an audio mixture based on a natural language query of the target source (e.g., "a man tells a joke followed by people laughing"). A unique challenge in LASS is associated with the complexity of natural language description and its relation with the audio sources. To address this issue, we proposed LASS-Net, an end-to-end neural network that is learned to jointly process acoustic and linguistic information, and separate the target source that is consistent with the language query from an audio mixture. We evaluate the performance of our proposed system with a dataset created from the AudioCaps dataset. Experimental results show that LASS-Net achieves considerable improvements over baseline methods. Furthermore, we observe that LASS-Net achieves promising generalization results when using diverse human-annotated descriptions as queries, indicating its potential use in real-world scenarios. The separated audio samples and source code are available at https://liuxubo717.github.io/LASS-demopage. Xubo Liu 0001, Haohe Liu, Qiuqiang Kong, Xinhao Mei, Jinzheng Zhao, Qiushi Huang, Mark D. Plumbley, Wenwu Wang 0001 |
INTERSPEECH | 5 |
| 2022 | Audio Visual Multi-Speaker Tracking with Improved GCF and PMBM Filter
Jinzheng Zhao, Peipei Wu, Xubo Liu 0001, Shidrokh Goudarzi, Haohe Liu, Yong Xu 0004, Wenwu Wang 0001 |
INTERSPEECH | 1 |
| 2019 | Robust Real-Time Object Detection Based on Deep Learning for Very High Resolution Remote Sensing ImagesabstractRecently, the development of deep learning boosts the object detection for remote sensing images. The existing deep learning methods can be divided into two types. The region-based methods represented by Faster R-CNN have progressive performance in accuracy. However, their computational cost is massive due to the deep Convolutional Neural Network (CNN) backbones, which limits the efficiency. The regression-based methods such as YOLO and Single Shot MultiBox Detector (SSD) are advantageous in speed while the accuracy is not satisfactory. To meet the increasing demand in both speed and accuracy for object detection of remote sensing images, we employ the Reception Field Block Net (RFBNet) detector. It embeds the Receptive Field Block (RFB) module into SSD to obtain better feature representation. The experimental results on NWPU VHR-10 dataset demonstrate that the mAP of RFBNet-512 reaches 91.56%, which outperforms other state-of-the-art networks. Meanwhile, the speed is also competitive. Jinzheng Zhao, Weiyu Xiong, Qingli Li, Junli Yang |
IGARSS | 2 |