Feijun Jiang

dblp:76/10697 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0001-5579-5144ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 ADC: Enhancing Function Calling Via Adversarial Datasets and Code Line-Level Feedback
abstract
Large Language Models (LLMs) have made significant strides in Natural Language Processing and coding, yet they struggle with robustness and accuracy in complex function calls. To tackle these challenges, this paper introduces ADC, an innovative approach that enhances LLMs’ ability to follow function formats and match complex parameters. ADC utilizes a high-quality code fine-tuning dataset with line-level execution feedback, providing granular process supervision that fosters strong logical reasoning and adherence to function formats. It also employs an adversarial dataset generation process to improve parameter matching. The staged training methodology capitalizes on both enriched code datasets and refined adversarial datasets, leading to marked improvements in function calling capabilities on the Berkeley Function-Calling Leaderboard (BFCL) Benchmark. The innovation of ADC lies in its strategic combination of process supervision, adversarial refinement, and incremental learning, setting a new standard for LLM proficiency in complex function calling.
Wei Zhang 0384, Qianghuai Jia, Feijun Jiang, Hongcheng Guo, Zhoujun Li 0001, Mengping Zhou
ICASSP5
2025 Sentence-graph-level knowledge injection with multi-task learning
Liyi Chen 0003, Yifei Yuan 0002, Jie Liu 0007, Feijun Jiang
World Wide Web (WWW)7
2024 CO3: Low-resource Contrastive Co-training for Generative Conversational Query Rewrite
abstract
Generative query rewrite generates reconstructed query rewrites using the conversation history while rely heavily on gold rewrite pairs that are expensive to obtain. Recently, few-shot learning is gaining increasing popularity for this task, whereas these methods are sensitive to the inherent noise due to limited data size. Besides, both attempts face performance degradation when there exists language style shift between training and testing cases. To this end, we study low-resource generative conversational query rewrite that is robust to both noise and language style shift. The core idea is to utilize massive unlabeled data to make further improvements via a contrastive co-training paradigm. Specifically, we co-train two dual models (namely Rewriter and Simplifier) such that each of them provides extra guidance through pseudo-labeling for enhancing the other in an iterative manner. We also leverage contrastive learning with data augmentation, which enables our model pay more attention on the truly valuable information than the noise. Extensive experiments demonstrate the superiority of our model under both few-shot and zero-shot scenarios. We also verify the better generalization ability of our model when encountering language style shift.
Yifei Yuan 0002, Liyi Chen 0003, Renjun Hu, Zengming Zhang, Feijun Jiang, Wai Lam
LREC/COLING7
2023 CHMATCH: Contrastive Hierarchical Matching and Robust Adaptive Threshold Boosted Semi-Supervised Learning
abstract
The recently proposed FixMatch and FlexMatch have achieved remarkable results in the field of semi-supervised learning. But these two methods go to two extremes as FixMatch and FlexMatch use a pre-defined constant threshold for all classes and an adaptive threshold for each category, respectively. By only investigating consistency regularization, they also suffer from unstable results and indiscriminative feature representation, especially under the situation of few labeled samples. In this paper, we propose a novel CHMatch method, which can learn robust adaptive thresholds for instance-level prediction matching as well as discriminative features by contrastive hierarchical matching. We first present a memory-bank based robust threshold learning strategy to select highly-confident samples. In the meantime, we make full use of the structured information in the hierarchical labels to learn an accurate affinity graph for contrastive learning. CHMatch achieves very stable and superior results on several commonly-used benchmarks. For example, CHMatch achieves 8.44% and 9.02% error rate reduction over FlexMatch on CIFAR-100 under WRN-28-2 with only 4 and 25 labeled samples per class, respectively11Project address: https://github.com/sailist/CHMatch.
Jianlong Wu, Haozhe Yang, Tian Gan 0002, Ning Ding 0006, Feijun Jiang, Liqiang Nie
CVPR5
2023 Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
abstract
In recent research, slight performance improvement is observed from automatic speech recognition systems to audio-visual speech recognition systems in end-to-end frameworks with low-quality videos. Unmatching convergence rates and specialized input representations between audio-visual modalities are considered to cause the problem. In this paper, we propose two novel techniques to improve audio-visual speech recognition (AVSR) under a pre-training and fine-tuning training framework. First, we explore the correlation between lip shapes and syllable-level subword units in Mandarin through a frame-level subword unit classification task with visual streams as input. The fine-grained subword labels guide the network to capture temporal relationships between lip shapes and result in an accurate alignment between video and audio streams. Next, we propose an audio-guided Cross-Modal Fusion Encoder (CMFE) to utilize main training parameters for multiple cross-modal attention layers to make full use of modality complementarity. Experiments on the MISP2021-AVSR data set show the effectiveness of the two proposed techniques. Together, using only a relatively small amount of training data, the final system achieves better performances than state-of-the-art systems with more complex front-ends and back-ends. The code is released at1.
Yusheng Dai, Hang Chen 0001, Jun Du 0002, Xiaofei Ding, Feijun Jiang, Chin-Hui Lee 0001
ICME6
2023 Bootstrap Latent Representations for Multi-modal Recommendation
abstract
This paper studies the multi-modal recommendation problem, where the item multi-modality information (e.g., images and textual descriptions) is exploited to improve the recommendation accuracy. Besides the user-item interaction graph, existing state-of-the-art methods usually use auxiliary graphs (e.g., user-user or item-item relation graph) to augment the learned representations of users and/or items. These representations are often propagated and aggregated on auxiliary graphs using graph convolutional networks, which can be prohibitively expensive in computation and memory, especially for large graphs. Moreover, existing multi-modal recommendation methods usually leverage randomly sampled negative examples in Bayesian Personalized Ranking (BPR) loss to guide the learning of user/item representations, which increases the computational cost on large graphs and may also bring noisy supervision signals into the training process. To tackle the above issues, we propose a novel self-supervised multi-modal recommendation model, dubbed BM3, which requires neither augmentations from auxiliary graphs nor negative samples. Specifically, BM3 first bootstraps latent contrastive views from the representations of users and items with a simple dropout augmentation. It then jointly optimizes three multi-modal objectives to learn the representations of users and items by reconstructing the user-item interaction graph and aligning modality features under both inter- and intra-modality perspectives. BM3 alleviates both the need for contrasting with negative examples and the complex graph augmentation from an additional target network for contrastive view generation. We show BM3 outperforms prior recommendation models on three datasets with number of nodes ranging from 20K to 200K, while achieving a 2-9 × reduction in training time. Code implementation is located at: https://github.com/enoche/BM3.
Xin Zhou 0008, Yong Liu 0020, Chunyan Miao, Pengwei Wang 0005, Yuan You, Feijun Jiang
WWW8
2023 iEmoTTS: Toward Robust Cross-Speaker Emotion Transfer and Control for Speech Synthesis Based on Disentanglement Between Prosody and Timbre
abstract
The capability of generating speech with a specific type of emotion is desired for many human-computer interaction applications. Cross-speaker emotion transfer is a common approach to generating emotional speech when speech data with emotion labels from target speakers is not available for model training. This paper presents a novel cross-speaker emotion transfer system named iEmoTTS. The system is composed of an emotion encoder, a prosody predictor, and a timbre encoder. The emotion encoder extracts the identity of emotion type and the respective emotion intensity from the mel-spectrogram of input speech. The emotion intensity is measured by the posterior probability that the input utterance carries that emotion. The prosody predictor is used to provide prosodic features for emotion transfer. The timbre encoder provides timbre-related information for the system. Unlike many other studies which focus on disentangling speaker and style factors of speech, the iEmoTTS is designed to achieve cross-speaker emotion transfer via disentanglement between prosody and timbre. Prosody is considered the primary carrier of emotion-related speech characteristics, and timbre accounts for the essential characteristics for speaker identification. Zero-shot emotion transfer, meaning that the speech of target speakers is not seen in model training, is also realized with iEmoTTS. Extensive experiments of subjective evaluation have been carried out. The results demonstrate the effectiveness of iEmoTTS compared with other recently proposed systems of cross-speaker emotion transfer. It is shown that iEmoTTS can produce speech with designated emotion types and controllable emotion intensity. With appropriate information bottleneck capacity, iEmoTTS is able to transfer emotional information to a new speaker effectively. Audio samples are publicly available.
Guangyan Zhang, Jialun Wu, Yutao Gai, Feijun Jiang, Tan Lee
IEEE ACM Trans. Audio Speech Lang. Process.7
2023 Neighbor-Guided Consistent and Contrastive Learning for Semi-Supervised Action Recognition
abstract
Semi-supervised learning has been well established in the area of image classification but remains to be explored in video-based action recognition. FixMatch is a state-of-the-art semi-supervised method for image classification, but it does not work well when transferred directly to the video domain since it only utilizes the single RGB modality, which contains insufficient motion information. Moreover, it only leverages highly-confident pseudo-labels to explore consistency between strongly-augmented and weakly-augmented samples, resulting in limited supervised signals, long training time, and insufficient feature discriminability. To address the above issues, we propose neighbor-guided consistent and contrastive learning (NCCL), which takes both RGB and temporal gradient (TG) as input and is based on the teacher-student framework. Due to the limitation of labelled samples, we first incorporate neighbors information as a self-supervised signal to explore the consistent property, which compensates for the lack of supervised signals and the shortcoming of long training time of FixMatch. To learn more discriminative feature representations, we further propose a novel neighbor-guided category-level contrastive learning term to minimize the intra-class distance and enlarge the inter-class distance. We conduct extensive experiments on four datasets to validate the effectiveness. Compared with the state-of-the-art methods, our proposed NCCL achieves superior performance with much lower computational cost.
Jianlong Wu, Tian Gan 0002, Ning Ding 0006, Feijun Jiang, Jialie Shen 0001, Liqiang Nie
IEEE Trans. Image Process.5
2022 McQueen: a Benchmark for Multimodal Conversational Query Rewrite
abstract
The task of query rewrite aims to convert an in-context query to its fully-specified version where ellipsis and coreference are completed and referred-back according to the history context.Although much progress has been made, less efforts have been paid to real scenario conversations that involve drawing information from more than one modalities.In this paper, we propose the task of multimodal conversational query rewrite (McQR), which performs query rewrite under the multimodal visual conversation setting.We collect a largescale dataset named McQueen based on manual annotation, which contains 15k visual conversations and over 80k queries where each one is associated with a fully-specified rewrite version.In addition, for entities appearing in the rewrite, we provide the corresponding image box annotation.We then use the McQueen dataset to benchmark a state-of-the-art method for effectively tackling the McQR task, which is based on a multimodal pre-trained model with pointer generator.Extensive experiments are performed to demonstrate the effectiveness of our model on this task 1 .
Yifei Yuan 0002, Liyi Chen 0003, Feijun Jiang, Yuan You, Wai Lam
EMNLP5
2022 Speech2Slot: A Limited Generation Framework with Boundary Detection for Slot Filling from Speech
Pengwei Wang 0005, Yinpei Su, Xiaohuan Zhou, Liangchen Wei, Yuan You, Feijun Jiang
INTERSPEECH8
2021 An End-to-End Far-Field Keyword Spotting System with Neural Beamforming
abstract
Conventional keyword spotting (KWS) systems typically use microphone array techniques to improve robustness against noise and reverberation. However, KWS systems with such traditional speech enhancement methods may not always yield the optimal performance in different conditions, as the optimization criterion for the speech enhancement part is not directly relevant to the KWS objective. In this work, we explore the KWS system by encompassing neural beamforming for speech enhancement within the KWS neural network, which updates both parameters by jointly optimizing the unique KWS criteria. We demonstrate that the proposed neural beamforming KWS system not only significantly outperforms the traditional KWS method by improving 30% relative recall rate at the same precision, but also can obviously reduce the system complexity, which could be much easier to be adopted by small resource required devices.
Fuming Fang, Dongdi Zhao, Feijun Jiang
ASRU9
2021 User Feedback and Ranking in-a-Loop: Towards Self-Adaptive Dialogue Systems
abstract
Accurate skill retrieval is a key factor for the success of modern conversational AI agents. The major challenges lie in the ambiguity in human spoken language and the wide spectrum of candidate skills. In this paper, we make the first attempt to attack the problem by implementing a user feedback enhanced reranking strategy, and propose a self-adaptive dialogue system (AdaDial) for conversational AI agents. In AdaDial, we consider estimating user feedback and adjusting ranking strategy into a "closed-loop". In particular, we propose a scalable schema for user feedback estimation and a feedback enhanced reranking model with customized feature encoding, target attention based feature assembling, and multi-task learning. As a result, AdaDial achieves self-adaptivity at both individual- and system-levels. Online experimental results demonstrate that AdaDial could not only retrieve desired skills for different users in different scenarios, but also correct its regular strategy according to negative feedback. AdaDial has been deployed on a large-scale conversational AI agents with tens of millions daily queries, and is bringing continued positive impacts on user experience.
Zengming Zhang, Feijun Jiang
SIGIR5
2021 Adapting to Context-Aware Knowledge in Natural Conversation for Multi-Turn Response Selection
abstract
Virtual assistants aim to build a human-like conversational agent. However, current human-machine conversations still cannot make users feel intelligent enough to build a continued dialog over time. Some responses from agents are usually inconsistent, uninformative, less-engaging and even memoryless. In recent years, most researchers have tried to employ conversation context and external knowledge, e.g. wiki pages and knowledge graphs, into the model which only focuses on solving some special conversation problems in local perspectives. Few researchers are dedicated to the whole capability of the conversational agent which is endowed with abilities of not only passively reacting the conversation but also proactively leading the conversation.
Chen Zhang 0003, Hao Wang 0005, Feijun Jiang, Hongzhi Yin
WWW3
2017 A Hybrid Framework for Text Modeling with Convolutional RNN
abstract
In this paper, we introduce a generic inference hybrid framework for Convolutional Recurrent Neural Network (conv-RNN) of semantic modeling of text, seamless integrating the merits on extracting different aspects of linguistic information from both convolutional and recurrent neural network structures and thus strengthening the semantic understanding power of the new framework. Besides, based on conv-RNN, we also propose a novel sentence classification model and an attention based answer selection model with strengthening power for the sentence matching and classification respectively. We validate the proposed models on a very wide variety of data sets, including two challenging tasks of answer selection (AS) and five benchmark datasets for sentence classification (SC). To the best of our knowledge, it is by far the most complete comparison results in both AS and SC. We empirically show superior performances of conv-RNN in these different challenging tasks and benchmark datasets and also summarize insights on the performances of other state-of-the-arts methodologies.
Feijun Jiang, Hongxia Yang
KDD2
2013 Combining texture and stereo disparity cues for real-time face detection
Feijun Jiang, Mika Fischer, Hazim Kemal Ekenel, Bertram E. Shi
Signal Process. Image Commun.1
2012 Efficient and robust integration of face detection and head pose estimation
Feijun Jiang, Hazim Kemal Ekenel, Bertram E. Shi
ICPR1
2011 Effective discretization of Gabor features for real-time face detection
abstract
We describe a real-time face detector based on Gabor features. While Gabor features often lead to improved performance, they are often avoided as they are perceived as being computationally expensive. We address this in two ways. First, we propose an efficient discrete encoding method for the Gabor feature vector. This enables us to use a computationally efficient multi-stage classifier based on boosting and winnowing. Second, we accelerate computationally complex computations using the parallelization provided by graphics processing units (GPUs). With these innovations, the resulting detector runs at 16.8 fps for 640 × 480 images on a PC equipped with an i5 CPU and a GTX 465 graphic card.
Feijun Jiang, Bertram E. Shi, Mika Fischer, Hazim Kemal Ekenel
ICIP1