Haojie Wei

dblp:236/5941 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 40% Question answering and dialogue systems · 21% Deep learning architectures and training · 13%
Computer graphics and multimedia
2 papers
Audio and music processing · 100%

Topics — the 14 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
music information retrieval
1.422024
MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch Estimation · ACM Multimedia 2024
JEPOO: Highly Accurate Joint Estimation of Pitch, Onset and Offset for Music Information Retrieval · IJCAI 2023
Natural language and speech › Language models and text generation › pre-trained language model › conversational language models
role-playing language models
0.912025
Beyond Dialogue: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model · ACL (1) 2025
Natural language and speech › Question answering and dialogue systems
question generation
0.822020
Asking Questions the Human Way: Scalable Question-Answer Generation from Text Corpus · WWW 2020
Learning to Generate Questions by LearningWhat not to Generate · WWW 2019
Machine learning › Deep learning architectures and training
mixture of experts
0.812024
Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs · EMNLP 2024
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.812024
Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs · EMNLP 2024
Audio and music processing › source separation
music source separation
0.812024
MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch Estimation · ACM Multimedia 2024
Audio and music processing › speech processing
pitch estimation
0.812024
MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch Estimation · ACM Multimedia 2024
Audio and music processing › music transcription
melody extraction
0.712023
JEPOO: Highly Accurate Joint Estimation of Pitch, Onset and Offset for Music Information Retrieval · IJCAI 2023
Natural language and speech › Question answering and dialogue systems › question generation
question-answer pair generation
0.412020
Asking Questions the Human Way: Scalable Question-Answer Generation from Text Corpus · WWW 2020
Machine learning › Learning paradigms
multi-task learning
0.422024
MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch Estimation · ACM Multimedia 2024
JEPOO: Highly Accurate Joint Estimation of Pitch, Onset and Offset for Music Information Retrieval · IJCAI 2023
Natural language and speech › Language models and text generation › text generation › neural text generation
copy mechanism
0.412019
Learning to Generate Questions by LearningWhat not to Generate · WWW 2019
Natural language and speech › Language models and text generation
text generation
0.412019
Learning to Generate Questions by LearningWhat not to Generate · WWW 2019
Machine learning › Optimization for machine learning › multi-objective optimization
pareto optimization
0.212023
JEPOO: Highly Accurate Joint Estimation of Pitch, Onset and Offset for Music Information Retrieval · IJCAI 2023
Natural language and speech › Information extraction and text analysis
textual entailment
0.112020
Asking Questions the Human Way: Scalable Question-Answer Generation from Text Corpus · WWW 2020

Methods — techniques the papers use, named apart from their topics

two-stage training · 1.5joint learning · 1.5dynamic weighting · 1.5pareto modulated loss · 1.3loss weight regularization · 1.3reasoning data generation · 0.9prompting · 0.9graph convolutional network · 0.8expert analysis · 0.8neural question generation · 0.4
YearPublicationVenuePosition
2025 Beyond Dialogue: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model
abstract
The rapid advancement of large language models (LLMs) has revolutionized role-playing, enabling the development of general role-playing models. However, current role-playing training has two significant issues: (I) Using a predefined role profile to prompt dialogue training for specific scenarios usually leads to biases and even conflicts between the dialogue and the profile, resulting in training biases. (II) Models learn to imitate the role based solely on the profile, neglecting profile-dialogue alignment at the sentence level. To overcome the aforementioned hurdles, we propose a novel framework Beyond Dialogue, which introduces “beyond dialogue” tasks to align dialogue with profile traits for each scenario, eliminating biases during training. Furthermore, the framework achieves a sentence-level fine-grained alignment between profile and dialogue through an innovative prompting mechanism that generates reasoning data for training. Moreover, the aforementioned methods are fully automated and low-cost. Experimental results demonstrate our model excels in adhering to role profiles, outperforming most proprietary general and specialized role-playing baselines. The code and data are provided in https://github.com/yuyouyu32/BeyondDialogue.
Yeyong Yu, Runsheng Yu, Haojie Wei, Zhanqiu Zhang, Quan Qian
ACL (1)3
2025 Incorporating prior knowledge for domain generalization traffic flow anomaly detection
Haojie Wei
Neural Comput. Appl.3
2025 Prototype-based dual-alignment of multi-source domain adaptation for radar emitter recognition
Haojie Wei, Haixiang Li, Zhanpeng Zheng
Signal Process.1
2025 PRADO: A Low-Latency and Energy-Efficient 6DoF Pose Refinement Accelerator With Domain-Specific Explorations
abstract
Six degrees of freedom (6DoF) pose estimation is a critical technique for applications involving humanoid robotics, autonomous driving, and virtual and augmented reality (VR/AR). Pose refinement plays a pivotal role in the 6DoF pose estimation, significantly enhancing accuracy by iteratively refining the initial pose derived from a single-shot pose estimation (SSPE) process. However, this iterative approach is time-and energy-consuming which is not suitable for resource-and power-constrained edge-devices. In this paper, we propose PRADO, an energy-efficient 6DoF pose refinement accelerator for edge applications. Several domain-specific explorations aimed at reducing processing latency and energy-consumption while maintaining high accuracy have been proposed, including an adaptive sparse correspondences sampling (ASCS) architecture to reduce redundant correspondences process, a hybrid static & dynamic pruning (HSDP) technique with co-designed hardware architecture to reduce computational complexity, and a cosine similarity-based adaptive early exit (CSAE2) technique to eliminate unnecessary iterations. Implemented and evaluated on the ZCU102 FPGA board, PRADO achieves the lowest processing latency of 61.2ms and energy consumption of 198.6mJ, and highest accuracy of 0.7112, outperforming state-of-the-art designs. Implemented with 28nm CMOS technology, PRADO achieves an even lower processing latency of 18.1ms and energy consumption of 12.6mJ.
Haojie Wei, Le Jin, Junhao Zeng, Liwei Zhuang, Yu Long 0005, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 An Energy-Efficient Block-Based Nonmaximum Suppression Engine for High-Parallel Postprocessing of Visual Object Detection
abstract
Nowadays, visual object detection (VOD) is widely used in many AI applications, such as autonomous driving, intelligent robotics, and smart surveillance. As an essential postprocessing step in VOD, nonmaximum suppression (NMS) is employed to generate bounding boxes as detection results. However, NMS is difficult to parallelize and computationally intensive, resulting in high processing latency and energy consumption. To address this issue, this brief proposes an energy-efficient block-based NMS engine that incorporates both algorithm- and hardware-level design techniques to improve processing speed and energy efficiency. These techniques include a block-NMS scheme, an adaptive hybrid sorting architecture (AHSA), and a reconfigurable pipeline-based block-NMS computation architecture. The proposed engine is implemented in 28-nm CMOS technology. Compared with the state-of-the-art designs, it achieves the highest performance (237.97 GOPS) and energy efficiency (8.71 TOPS/W), while delivering results fully equivalent to those of the original NMS.
Yuchuan Gong, Haojie Wei, Hongtao Guo, Jiahao Zheng 0004, Qingyuan Hou, Zherong Liu, Jingxiao Zheng, Ye Liu 0011, Zhengning Wang, Jun Zhou 0017
IEEE Trans. Very Large Scale Integr. Syst.2
2024 MIPM: A Multidimensional Information Perception Model for Estimating Time of Arrival on Real Road Networks
Tangpeng Dan, Yingtao Peng, Haojie Wei, Xiaofeng Meng 0001
DASFAA (1)4
2024 Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs
abstract
Xin Zhou, Ping Nie, Yiwen Guo, Haojie Wei, Zhanqiu Zhang, Pasquale Minervini, Ruotian Ma, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Xin Zhou 0012, Ping Nie, Yiwen Guo, Haojie Wei, Zhanqiu Zhang, Pasquale Minervini, Ruotian Ma, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP4
2024 DJCM: A Deep Joint Cascade Model for Singing Voice Separation and Vocal Pitch Estimation
abstract
Singing voice separation and vocal pitch estimation are pivotal tasks in music information retrieval. Existing methods for simultaneous extraction of clean vocals and vocal pitches can be classified into two categories: pipeline methods and naive joint learning methods. However, the efficacy of these methods is limited by the following problems: On the one hand, pipeline methods train models for each task independently, resulting a mismatch between the data distributions at the training and testing time. On the other hand, naive joint learning methods simply add the losses of both tasks, possibly leading to a misalignment between the distinct objectives of each task. To solve these problems, we propose a Deep Joint Cascade Model (DJCM) for singing voice separation and vocal pitch estimation. DJCM employs a novel joint cascade model structure to concurrently train both tasks. Moreover, task-specific weights are used to align different objectives of both tasks. Experimental results show that DJCM achieves state-of-the-art performance on both tasks, with great improvements of 0.45 in terms of Signal-to-Distortion Ratio (SDR) for singing voice separation and 2.86% in terms of Overall Accuracy (OA) for vocal pitch estimation. Furthermore, extensive ablation studies validate the effectiveness of each design of our proposed model.
Haojie Wei, Xueke Cao, Tangpeng Dan, Yueguo Chen
ICASSP1
2024 MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch Estimation
abstract
Music source separation and pitch estimation are two vital tasks in music information retrieval. Typically, the input of pitch estimation is obtained from the output of music source separation. Therefore, existing methods have tried to perform these two tasks simultaneously, so as to leverage the mutually beneficial relationship between both tasks. However, these methods still face two critical challenges that limit the improvement of both tasks: the lack of labeled data and joint learning optimization. To address these challenges, we propose a Model-Agnostic Joint Learning (MAJL) framework for both tasks. MAJL is a generic framework and can use variant models for each task. It includes a two-stage training method and a dynamic weighting method named Dynamic Weights on Hard Samples (DWHS), which addresses the lack of labeled data and joint learning optimization, respectively. Experimental results on public music datasets show that MAJL outperforms state-of-the-art methods on both tasks, with significant improvements of 0.92 in Signal-to-Distortion Ratio (SDR) for music source separation and 2.71% in Raw Pitch Accuracy (RPA) for pitch estimation. Furthermore, comprehensive studies not only validate the effectiveness of each component of MAJL, but also indicate the great generality of MAJL in adapting to different model architectures.
Haojie Wei, Jun Yuan 0008, Rui Zhang 0003, Quanyu Dai, Yueguo Chen
ACM Multimedia1
2023 JEPOO: Highly Accurate Joint Estimation of Pitch, Onset and Offset for Music Information Retrieval
abstract
Melody extraction is a core task in music information retrieval, and the estimation of pitch, onset and offset are key sub-tasks in melody extraction. Existing methods have limited accuracy, and work for only one type of data, either single-pitch or multi-pitch. In this paper, we propose a highly accurate method for joint estimation of pitch, onset and offset, named JEPOO. We address the challenges of joint learning optimization and handling both single-pitch and multi-pitch data through novel model design and a new optimization technique named Pareto modulated loss with loss weight regularization. This is the first method that can accurately handle both single-pitch and multi-pitch music data, and even a mix of them. A comprehensive experimental study on a wide range of real datasets shows that JEPOO outperforms state-of-the-art methods by up to 10.6\%, 8.3\% and 10.3\% for the prediction of Pitch, Onset and Offset, respectively, and JEPOO is robust for various types of data and instruments. The ablation study validates the effectiveness of each component of JEPOO.
Haojie Wei, Jun Yuan 0008, Rui Zhang 0003, Yueguo Chen
IJCAI1
2023 RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
abstract
Vocal pitch is an important high-level feature in music audio processing.However, extracting vocal pitch in polyphonic music is more challenging due to the presence of accompaniment.To eliminate the influence of the accompaniment, most previous methods adopt music source separation models to obtain clean vocals from polyphonic music before predicting vocal pitches.As a result, the performance of vocal pitch estimation is affected by the music source separation models.To address this issue and directly extract vocal pitches from polyphonic music, we propose a robust model named RMVPE.This model can extract effective hidden features and accurately predict vocal pitches from polyphonic music.The experimental results demonstrate the superiority of RMVPE in terms of raw pitch accuracy (RPA) and raw chroma accuracy (RCA).Additionally, experiments conducted with different types of noise show that RMVPE is robust across all signal-to-noise ratio (SNR) levels.The code of RMVPE is available at https: //github.com/Dream-High/RMVPE.
Haojie Wei, Xueke Cao, Tangpeng Dan, Yueguo Chen
INTERSPEECH1
2020 Asking Questions the Human Way: Scalable Question-Answer Generation from Text Corpus
abstract
The ability to ask questions is important in both human and machine intelligence. Learning to ask questions helps knowledge acquisition, improves question-answering and machine reading comprehension tasks, and helps a chatbot to keep the conversation flowing with a human. Existing question generation models are ineffective at generating a large amount of high-quality question-answer pairs from unstructured text, since given an answer and an input passage, question generation is inherently a one-to-many mapping. In this paper, we propose Answer-Clue-Style-aware Question Generation (ACS-QG), which aims at automatically generating high-quality and diverse question-answer pairs from unlabeled text corpus at scale by imitating the way a human asks questions. Our system consists of: i) an information extractor, which samples from the text multiple types of assistive information to guide question generation; ii) neural question generators, which generate diverse and controllable questions, leveraging the extracted assistive information; and iii) a neural quality controller, which removes low-quality generated data based on text entailment. We compare our question generation models with existing approaches and resort to voluntary human evaluation to assess the quality of the generated question-answer pairs. The evaluation results suggest that our system dramatically outperforms state-of-the-art neural question generation models in terms of the generation quality, while being scalable in the meantime. With models trained on a relatively smaller amount of data, we can generate 2.8 million quality-assured question-answer pairs from a million sentences found in Wikipedia.
Bang Liu 0003, Haojie Wei, Di Niu 0002, Haolan Chen, Yancheng He
WWW2
2019 Matching Article Pairs with Graphical Decomposition and Convolutions
abstract
Identifying the relationship between two articles, e.g., whether two articles published from different sources describe the same breaking news, is critical to many document understanding tasks.Existing approaches for modeling and matching sentence pairs do not perform well in matching longer documents, which embody more complex interactions between the enclosed entities than a sentence does.To model article pairs, we propose the Concept Interaction Graph to represent an article as a graph of concepts.We then match a pair of articles by comparing the sentences that enclose the same concept vertex through a series of encoding techniques, and aggregate the matching signals through a graph convolutional network.To facilitate the evaluation of long article matching, we have created two datasets, each consisting of about 30K pairs of breaking news articles covering diverse topics in the open domain.Extensive evaluations of the proposed methods on the two datasets demonstrate significant improvements over a wide range of state-of-the-art methods for natural language matching.
Bang Liu 0003, Di Niu 0002, Haojie Wei, Jinghong Lin, Yancheng He, Kunfeng Lai
ACL (1)3
2019 Learning to Generate Questions by LearningWhat not to Generate
abstract
Automatic question generation is an important technique that can improve the training of question answering, help chatbots to start or continue a conversation with humans, and provide assessment materials for educational purposes. Existing neural question generation models are not sufficient mainly due to their inability to properly model the process of how each word in the question is selected, i.e., whether repeating the given passage or being generated from a vocabulary. In this paper, we propose our Clue Guided Copy Network for Question Generation (CGC-QG), which is a sequence-to-sequence generative model with copying mechanism, yet employing a variety of novel components and techniques to boost the performance of question generation. In CGC-QG, we design a multi-task labeling strategy to identify whether a question word should be copied from the input passage or be generated instead, guiding the model to learn the accurate boundaries between copying and generation. Furthermore, our input passage encoder takes as input, among a diverse range of other features, the prediction made by a clue word predictor, which helps identify whether each word in the input passage is a potential clue to be copied into the target question. The clue word predictor is designed based on a novel application of Graph Convolutional Networks onto a syntactic dependency tree representation of each passage, thus being able to predict clue words only based on their context in the passage and their relative positions to the answer in the tree. We jointly train the clue prediction as well as question generation with multi-task learning and a number of practical strategies to reduce the complexity. Extensive evaluations show that our model significantly improves the performance of question generation and out-performs all previous state-of-the-art neural question generation models by a substantial margin.
Bang Liu 0003, Mingjun Zhao, Di Niu 0002, Kunfeng Lai, Yancheng He, Haojie Wei
WWW6