Feiyu Zhao

dblp:231/1622 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-5747-7614ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
abstract
Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks.However, hallucination, where models generate responses that are semantically incorrect or acoustically unsupported, remains largely underexplored in the audio domain.Existing hallucination benchmarks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth.We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music.HalluAudio comprises over 5K humanverified QA pairs and spans diverse task types, including binary judgments, multi-choice reasoning, attribute verification, and open-ended QA.To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions.Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, enabling a fine-grained analysis of LALM failure modes.We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music.Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs.
Feiyu Zhao, Wenhuan Lu, Daipeng Zhang 0001, Xianghu Yue, Jianguo Wei
ACL (1)1
2026 SegRNN: Segment Recurrent Neural Network for Long-Term Time-Series Forecasting
abstract
With the proliferation of Internet of Things (IoT) applications, advanced time series forecasting techniques have become increasingly critical for managing and responding to complex temporal dynamics. However, traditional RNN-based methods have faced challenges in the Long-term Time Series Forecasting (LTSF) domain when dealing with excessively long look-back windows and forecast horizons. Consequently, the dominance in this domain has shifted towards Transformer, MLP, and CNN approaches. The substantial number of recurrent iterations are the fundamental reasons behind the limitations of RNNs in LTSF. To address these issues, we propose two novel strategies to reduce the number of iterations in RNNs for LTSF tasks: Segment-wise Iterations and Parallel Multi-step Forecasting (PMF). RNNs that combine these strategies, called SegRNN, significantly reduce the required recurrent iterations for LTSF, resulting in notable improvements in forecast accuracy and inference speed. Extensive experiments demonstrate that SegRNN not only outperforms state-of-the-art Transformer-based models but also reduces runtime and memory usage by more than 78%, making it highly suitable for resource-constrained IoT scenarios. These achievements provide strong evidence that RNNs continue to excel in LTSF tasks and encourage further exploration of this domain with more RNN-based approaches. The code is available at: https://github.com/lss-1138/SegRNN.
Shengsheng Lin, Weiwei Lin 0001, Wentai Wu, Feiyu Zhao, Ruichao Mo, Haotong Zhang 0003
IEEE Internet Things J.4
2026 OF-DETR: an efficient end-to-end detector for tiny traffic targets in aerial images
Hanzhang Huang, Feiyu Zhao, Yuxuan Tang, Shuaidi He, Xinghao Chen 0011, Qixiang Guo, Minchao Zhang
Multim. Syst.3
2025 Initiating a novel elementary school artificial intelligence-related image recognition curricula
Feiyu Zhao, Yongming Chen
Multim. Tools Appl.2
2025 MSCNet: Multi-Scale Network With Convolutions for Long-Term Cloud Workload Prediction
abstract
Accurate workload prediction is crucial for resource allocation and management in large-scale cloud data centers. While many approaches have been proposed, most existing methods are based on Recurrent Neural Networks (RNNs) or their variants, focusing on short-term cloud workload prediction without considering or identifying the long-term changes and different periodic patterns of cloud workloads. Due to variations in user demands or workload dynamics, cloud workloads that appear stable in the short term often exhibit distinct patterns in the long term. This can lead to a significant decline in prediction accuracy for existing methods when applied to long-term cloud workload forecasting. To address these challenges and overcome the limitations of current approaches, we propose a Multi-Scale Network with Convolutions (MSCNet) for accurate long-term cloud workload prediction. MSCNet employs multi-scale modeling of the original cloud workload to effectively extract multi-scale features and different periodic patterns, learning the long-term dependencies among the cloud workload. Our core component, the Multi-Scale Block, combines the Multi-Scale Patch Block, Transformer Encoder, and Multi-Scale Convolutions Block for comprehensive multi-scale learning. This enables MSCNet to adaptively learn both short-term and long-term features and patterns of cloud workloads, resulting in accurate long-term cloud workload predictions. Extensive experiments are conducted using real-world cloud workload data from Alibaba, Google, and Azure to validate the effectiveness of MSCNet. The experimental results demonstrate that MSCNet achieves accurate long-term cloud workload prediction with a computational complexity of$O(L^{2}d)$, outperforming existing state-of-the-art methods.
Feiyu Zhao, Weiwei Lin 0001, Shengsheng Lin, Shaomin Tang, Keqin Li 0001
IEEE Trans. Serv. Comput.1
2025 TFEGRU: Time-Frequency Enhanced Gated Recurrent Unit With Attention for Cloud Workload Prediction
abstract
Accurate prediction of cloud workload is crucial for effective resource allocation in cloud computing. However, due to the complexity and high dimensionality of workloads in the cloud environment, achieving precise workload prediction is a complex and challenging problem. Current approaches to cloud workload prediction mainly rely on deep learning methods based on the Recurrent Neural Network (RNN), which struggle to capture the long-term dependencies inherent in workloads effectively. To tackle these challenges and overcome the limitations of existing methods, we propose an effective approach Time-Frequency Enhanced Gated Recurrent Unit with Attention (TFEGRU) for cloud workload prediction. First, we design a Time-Frequency Enhanced Block (TFEB) to capture complex workload patterns and extract features from both the frequency and temporal domains. Next, we integrate channel independent strategy and channel embedding into the model to adapt to high-dimensional workloads and enhance predictive performance. Finally, we apply a Gated Recurrent Unit (GRU) in conjunction with a multi-head self-attention mechanism to achieve accurate workload prediction. To validate the effectiveness of TFEGRU, comprehensive experiments are conducted using real-world traces from Google and Alibaba cloud data centers. The experimental results demonstrate that TFEGRU achieves accurate and efficient predictions across diverse cloud workloads, outperforming existing state-of-the-art methods.
Feiyu Zhao, Weiwei Lin 0001, Shengsheng Lin, Haocheng Zhong, Keqin Li 0001
IEEE Trans. Serv. Comput.1
2024 Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration
abstract
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x.github.io.
Abigail O'Neill, Abhiram Maddukuri, Abhishek Gupta 0004, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew E. Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu 0003, Charlotte Le, Chelsea Finn, Chen Wang 0053, Chenfeng Xu, Cheng Chi 0001, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Drieß, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Paul Foster, Fangchen Liu, Federico Ceola, Fei Xia 0002, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su 0001, Haoshu Fang, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim 0001, Jaimyn Drake, Jan Peters 0001, Jan Schneider 0007, Jasmine Hsu, Jeannette Bohg, Jeffrey T. Bingham, Jensen Gao, Jiaheng Hu, Jiajun Wu 0001, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan 0001, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Kenneth Y. Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Zhang 0002, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Yunliang Chen 0001, Lerrel Pinto, Li Fei-Fei 0001, Liam Tan, Linxi Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang 0003, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma 0001, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen 0001, Nicolas Heess, Nikhil J. Joshi, Niko Sünderhauf, Norman Di Palo, Nur Muhammad Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R. Sanketi, Patrick Tree Miller, Patrick Yin, Paul Wohlhart, Peng Xu 0010, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Rafael Rafailov, Ria Doshi, Roberto Martin Martin, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante-Gomez, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham D. Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair 0003, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wolfram Burgard, Xiaolong Wang 0004, Xinghao Zhu, Xinyang Geng, Liangwei Xu, Yecheng Jason Ma 0001, Yejin Kim 0003, Yevgen Chebotar, Yilin Wu 0003, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang 0001, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zichen Jeff Cui, Zichen Zhang 0016, Zipeng Lin
ICRA63
2024 End-To-End Speaker Anonymization Based on Location-Variable Convolution and Multi-Head Self-Attention
abstract
Speaker anonymization, a user-centric solution for voice privacy, aims to conceal the speaker’s identity while maintaining clarity and naturalness. The prevalent approach involves cascading modules of automatic speech recognition (ASR) and text-to-speech (TTS) models for speaker anonymization through speech synthesis. However, the inherent multimodal cascade nature of this approach leads to high error rates and unclear speech due to inaccuracies propagated from the ASR system to the TTS system. To address these issues, this paper proposes an end-to-end method for achieving zero-shot speaker anonymization. This method improves the shortcomings of traditional speech models that use a large number of fixed convolution kernels to capture the internal dependencies of speech sequences. It captures the internal dependencies of speech from both long-time and long-distance perspectives by combining a network with variable kernels, namely location-variable convolutions (LVCs), with a multi-head self-attention mechanism and dynamically adjusting weights. Besides, it learns anonymized speaker features flexibly through an enhanced cycle-consistency loss, iteratively aligning speaker information for restructured speech with that of anonymized speakers indefinitely. The efficacy of our proposed speaker anonymization model was demonstrated on the English dataset VCTK.
Feiyu Zhao, Jianguo Wei, Wenhuan Lu
TrustCom1
2024 Enhancing MeshNet for 3D shape classification with focal and regularization losses
Feiyu Zhao
Comput. Graph.2
2023 Comprehensive analysis of the heterogeneous computing performance of DNNs under typical frameworks on cloud and edge computing platforms
Feiyu Zhao, Yongming Chen
Expert Syst. Appl.1
2021 A novel multi-objective optimization algorithm for the integrated scheduling of flexible job shops considering preventive maintenance activities and transportation processes
Hui Wang 0055, Buyun Sheng, Qibing Lu, Xiyan Yin, Feiyu Zhao, Xincheng Lu, Ruiping Luo, Gaocai Fu
Soft Comput.5
2019 A novel heuristic algorithm with activity back-shift response model for resource-constrained project scheduling problem
Buyun Sheng, Hui Wang 0055, Chenglei Zhang, Feiyu Zhao, Xiyan Yin
Soft Comput.5
2018 An effective application of 3D cloud printing service quality evaluation in BM-MOPSO
abstract
Summary Addressing service control factors, rapid manufacturing environment change, difficulty of resource allocation evaluation, resource optimization of 3D cloud printing service in a cloud manufacturing environment, and other characteristics, this paper proposes an evaluation indicator system of innovative new product development 3D printing order task execution. The evaluation indicator has eight dimensional components, including Time (T), Quality of Service (Q), Matching (Mat), Reliability (R), Flexibility (Flex), Cost (C), Fault tolerance (Ft), and Satisfaction (Sa). It constructs a type of optimal selection model based on a Multi‐Agent 3D Cloud Printing Service Quality Evaluation and a framework of cloud service evaluation of an AHP‐TOPSIS evaluation model based on Pareto optimization, and it designs an algorithm involving hybrid multi‐objective particle swarm optimization (PSO) based on the Baldwin Effect Model. In addition, this paper verifies the effectiveness of the algorithm through an example and offers a case study designed to test its feasibility and effectiveness.
Xinggang Wang, Buyun Sheng, Chenglei Zhang, Hui Wang 0055, Feiyu Zhao
Concurr. Comput. Pract. Exp.6