Jinchao Zhang 0001

dblp:127/3143-1 · DBLP profile ↗
← Back
45ranked-venue papers
2as first author
30since 2021 · last 2026
0000-0003-4611-9675ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 2 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 19 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
abstract
Traditional dialogue retrieval aims to select the most appropriate utterance or image from recent dialogue history. However, they often fail to meet users’ actual needs for revisiting semantically coherent content scattered across long-form conversations. To fill this gap, we define the Fine-grained Fragment Retrieval (FFR) task, requiring models to locate query-relevant fragments, comprising both utterances and images, from multimodal long-form dialogues. As a foundation for FFR, we construct MLDR, the longest-turn multimodal dialogue retrieval dataset to date, averaging 25.45 turns per dialogue, with each naturally spanning three distinct topics. To evaluate generalization in real-world scenarios, we curate and annotate a WeChat-based test set comprising real-world multimodal dialogues with an average of 75.38 turns. Building on these resources, we explore existing generation-based Vision-Language Models (VLMs) on FFR and observe that they often retrieve incoherent utterance-image fragments. While optimized for generating responses from visual-textual inputs, these models lack explicit supervision to ensure semantic coherence within retrieved fragments. To address this, we propose F2RVLM, a generative retrieval model trained in a two-stage paradigm: (1) supervised fine-tuning to inject fragment-level retrieval knowledge, and (2) GRPO-based reinforcement learning with multi-objective rewards to encourage outputs with semantic precision, relevance, and contextual coherence. In addition, to account for difficulty variations arising from differences in intra-fragment element distribution, ranging from locally dense to sparsely scattered, we introduce a difficulty-aware curriculum sampling that ranks training instances by predicted difficulty and gradually incorporates harder examples. This strategy enhances the model’s reasoning ability in long-form, multi-turn dialogue contexts. Experiments on both in-domain and real-domain sets demonstrate that F2RVLM substantially outperforms popular VLMs, achieving superior retrieval performance.
Hanbo Bi, Zexi Jia, Jiapei Zhang, Peixiang Luo, Xiaoyue Duan, Jinchao Zhang 0001
AAAI9
2025 Investigating Context-Aware Collaborative Text Entry on Smartphones using Large Language Models
abstract
Text entry is a fundamental and ubiquitous task, but users often face challenges such as situational impairments or difficulties in sentence formulation.Motivated by this, we explore the potential of large language models (LLMs) to assist with text entry in realworld contexts.We propose a collaborative smartphone-based text entry system, CATIA, that leverages LLMs to provide text suggestions based on contextual factors, including screen content, time, location, activity, and more.In a 7-day in-the-wild study with 36 participants, the system offered appropriate text suggestions in over 80% of cases.Users exhibited different collaborative behaviors depending on whether they were composing text for interpersonal communication or information services.Additionally, the relevance
Yuanchun Shi, Weinan Shi, Meizhu Chen, Yeshuang Zhu, Jinchao Zhang 0001, Chun Yu
CHI9
2025 Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution Analysis
abstract
The advancement of Generative Adversarial Networks (GANs) and diffusion models significantly enhances the realism of synthetic images, driving progress in image processing and creative design. However, this progress also necessitates the development of effective detection methods, as synthetic images become increasingly difficult to distinguish from real ones. This difficulty leads to societal issues, such as the spread of misinformation, identity theft, and online fraud. While previous detection methods perform well on public benchmarks, they struggle with our benchmark, FakeART, particularly when dealing with the latest models and cross-domain tasks (e.g., photo-to-painting). To address this challenge, we develop a new synthetic image detection technique based on color distribution. Unlike real images, synthetic images often exhibit uneven color distribution. By employing color quantization and restoration techniques, we analyze the color differences before and after image restoration. We discover and prove that these differences closely relate to the uniformity of color distribution. Based on this finding, we extract effective color features and combine them with image features to create a detection model with only 1.4 million parameters. This model achieves state-of-the-art results across various evaluation benchmarks, including the challenging FakeART dataset.
Zexi Jia, Chuanwei Huang, Yeshuang Zhu, Hongyan Fei, Xiaoyue Duan, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016
CVPR9
2025 Semantic to Structure: Learning Structural Representations for Infringement Detection
abstract
Structural information in images is crucial for aesthetic assessment, and it is widely recognized in the artistic field that imitating the structure of other works significantly infringes on creators’ rights. The advancement of diffusion models has led to AI-generated content imitating artists’ structural creations, yet effective detection methods are still lacking. In this paper, we define this phenomenon as "structural infringement" and propose a corresponding detection method. Additionally, we develop quantitative metrics and create manually annotated datasets for evaluation: the SIA dataset of synthesized data, and the SIR dataset of real data. Due to the current lack of datasets for structural infringement detection, we propose a new data synthesis strategy based on diffusion models and LLM, successfully training a structural infringement detection model. Experimental results show that our method can successfully detect structural infringements and achieve notable improvements on annotated test sets.
Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jinchao Zhang 0001, Jie Zhou 0016
ICASSP6
2025 ILDiff: Generate Transparent Animated Stickers by Implicit Layout Distillation
abstract
High-quality animated stickers usually contain transparent channels, which are often ignored by current video generation models. To generate fine-grained animated transparency channels, existing methods can be roughly divided into video matting algorithms and diffusion-based algorithms. The methods based on video matting have poor performance in dealing with semi-open areas in stickers, while diffusion-based methods are often used to model a single image, which will lead to local flicker when modeling animated stickers. In this paper, we firstly propose an ILDiff method to generate animated transparent channels through implicit layout distillation, which solves the problems of semi-open area collapse and no consideration of temporal information in existing methods. Secondly, we create the Transparent Animated Sticker Dataset (TASD), which contains 0.32M high-quality samples with transparent channel, to provide data support for related fields. Extensive experiments demonstrate that ILDiff can produce finer and smoother transparent channels compared to other methods such as Matting Anything and Layer Diffusion. Our code and dataset will be released at link https://xiaoyuan1996.github.io.
Yeshuang Zhu, Jie Zhou 0016, Jinchao Zhang 0001
ICASSP5
2025 MCID: Multi-aspect Copyright Infringement Detection for Generated Images
Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jiapei Zhang, Xiaoyue Duan, Jinchao Zhang 0001, Jie Zhou 0016
ICCV9
2025 A Visual Leap in Clip Compositionality Reasoning Through Generation of Counterfactual Sets
abstract
Vision-language models (VLMs) often struggle with compositional reasoning due to insufficient high-quality image-text data. To tackle this challenge, we propose a novel block-based diffusion approach that automatically generates counterfactual datasets without manual annotation. Our method utilizes large language models to identify entities and their spatial relationships. It then independently generates image blocks as "puzzle pieces" coherently arranged according to specified compositional rules. This process creates diverse, high-fidelity counterfactual image-text pairs with precisely controlled variations. In addition, we introduce a specialized loss function that differentiates inter-set from intra-set samples, enhancing training efficiency and reducing the need for negative samples. Experiments demonstrate that fine-tuning VLMs with our counterfactual datasets significantly improves visual reasoning performance. Our approach achieves state-of-the-art results across multiple benchmarks while using substantially less training data than existing methods.
Zexi Jia, Chuanwei Huang, Hongyan Fei, Yeshuang Zhu, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016
ICCV8
2025 From Imitation to Innovation: The Emergence of Ai's Unique Artistic Styles and the Challenge of Copyright Protection
Zexi Jia, Chuanwei Huang, Yeshuang Zhu, Hongyan Fei, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016
ICCV8
2025 WalkVLM: Aid Visually Impaired People Walking by Vision Language Model
abstract
Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people. With the recent progress of vision-language models (VLMs), employing VLMs to improve this field has emerged as a popular research topic. However, most existing methods are studied on self-built question-answering datasets, lacking a unified training and testing benchmark for walk guidance. Moreover, in blind walking task, it is necessary to perform real-time streaming video parsing and generate concise yet informative reminders, which poses a great challenge for VLMs that suffer from redundant responses and low inference efficiency. In this paper, we firstly release a diverse, extensive, and unbiased walking awareness dataset, containing 12k video-manual annotation pairs from Europe and Asia to provide a fair training and testing benchmark for blind walking task. Furthermore, a WalkVLM model is proposed, which employs chain of thought for hierarchical planning to generate concise but informative reminders and utilizes temporal-aware adaptive prediction to reduce the temporal redundancy of reminders. Finally, we have established a solid benchmark for blind walking task and verified the advantages of WalkVLM in stream video processing for this task compared to other VLMs. Our dataset and code will be released at anonymous link https://walkvlm2024.github.io.
Yeshuang Zhu, Jiapei Zhang, Zexi Jia, Peixiang Luo, Xiaoyue Duan, Jie Zhou 0016, Jinchao Zhang 0001
ICCV10
2025 VSD2M: Large-scale Vision-language Sticker Dataset for Multi-frame Animated Sticker Generation
abstract
media, Nowadays, advanced text-to-video algorithms have spawned numerous general video generation systems that allow users to customize high-quality, photo-realistic videos by only providing simple text prompts. However, creating customized animated stickers, which have lower frame rates and more abstract semantics than videos, is greatly hindered by difficulties in data acquisition and incomplete benchmarks. To facilitate the exploration of researchers in animated sticker generation (ASG) field, we construct the currently largest vision-language sticker dataset named "VSD2M" at a two-million scale that contains static and animated stickers. Furthermore, to improve the performance of traditional video generation methods on ASG tasks with discrete characteristics, we propose a Spatial Temporal Interaction layer that utilizes semantic interaction and detail preservation to address the issue of insufficient information utilization. To our knowledge, this is the most comprehensive benchmark for multi-frame ASG task, and we hope it can provide valuable inspiration for other scholars in intelligent creation.
Jiapei Zhang, Yeshuang Zhu, Jie Zhou 0016, Jinchao Zhang 0001
ICME6
2025 A Synthetic-to-Real Dehazing Method based on Domain Unification
abstract
Due to distribution shift, the performance of deep learning-based method for image dehazing is adversely affected when applied to real-world hazy images. In this paper, we find that such deviation in dehazing task between real and synthetic domains may come from the imperfect collection of clean data. Owing to the complexity of the scene and the effect of depth, the collected clean data cannot strictly meet the ideal conditions, which makes the atmospheric physics model in the real domain inconsistent with that in the synthetic domain. For this reason, we come up with a synthetic-to-real dehazing method based on domain unification, which attempts to unify the relationship between the real and synthetic domain, thus to let the dehazing model more in line with the actual situation. Extensive experiments qualitatively and quantitatively demonstrate that the proposed dehazing method significantly outperforms state-of-the-art methods on real-world images.
Jie Zhou 0016, Jinchao Zhang 0001
ICME3
2025 MedDiT: A Knowledge-Controlled Diffusion Transformer Framework for Dynamic Medical Image Generation in Virtual Simulated Patient
abstract
Medical education relies heavily on Simulated Patients (SPs) to provide a safe environment for students to practice clinical skills, including medical image analysis. However, the high cost of recruiting qualified SPs and the lack of diverse medical imaging datasets have presented significant challenges. To address these issues, this paper introduces MedDiT, a novel knowledge-controlled conversational framework that can dynamically generate plausible medical images aligned with simulated patient symptoms, enabling diverse diagnostic skill training. Specifically, MedDiT integrates various patient Knowledge Graphs (KGs), which describe the attributes and symptoms of patients, to dynamically prompt Large Language Models' (LLMs) behavior and control the patient characteristics, mitigating hallucination during medical conversation. Additionally, a well-tuned Diffusion Transformer (DiT) model is incorporated to generate medical images according to the specified patient attributes in the KG. In this paper, we present the capabilities of MedDiT through a practical demonstration, showcasing its ability to act in diverse simulated patient cases and generate the corresponding medical images. This can provide an abundant and interactive learning experience for students, advancing medical education by offering an immersive simulation platform for future healthcare professionals. The work sheds light on the feasibility of incorporating advanced technologies like LLM, KG, and DiT in education applications, highlighting their potential to address the challenges faced in simulated patient-based medical education.
Yanzeng Li, Jinchao Zhang 0001, Jie Zhou 0016, Lei Zou 0001
IJCAI3
2025 Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) usually generate texts which satisfy context coherence but don't match the visual input. Such a hallucination issue hinders LVLMs' applicability in the real world. The key to solving hallucination in LVLM is to make the text generation rely more on the visual content. Most previous works choose to enhance/adjust the features/output of a specific modality (i.e., visual or textual) to alleviate hallucinations in LVLM, which do not explicitly or systematically enhance the visual reliance. In this paper, we comprehensively investigate the factors that may degenerate the visual reliance in text generation of LVLM from a Bayesian perspective. We propose to mitigate hallucination in LVLM from three aspects. Firstly, we observe that not all visual tokens are informative in generating meaningful texts. We propose to evaluate and remove redundant visual tokens to avoid their disturbance. Secondly, LVLM may encode inappropriate prior information, making it lean toward generating unexpected words. We propose a simple, yet effective way to rectify the prior from a Bayesian perspective. Thirdly, we observe that starting from certain steps, the posterior of next-token prediction conditioned on visual tokens may collapse to a prior distribution which does not depend on any informative visual tokens at all. Thus, we propose to stop further text generation to avoid hallucination. Extensive experiments on three benchmarks, including POPE, CHAIR, and MME, demonstrate that our method can consistently mitigate the hallucination issue of LVLM and performs favorably against previous state-of-the-arts. Codes are available at https://github.com/NeilHnxTcc/EVRB.
Nanxing Hu, Xiaoyue Duan, Jinchao Zhang 0001, Guoliang Kang
ACM Multimedia3
2025 ArtFRD: A Fisher-Rao Mixture Metric for Generative Model Aesthetic Evaluation
abstract
Recent advances in generative modeling have enabled the synthesis of high-quality artistic images. Nevertheless, systematic evaluation of generative models from an aesthetic standpoint is still lacking, which hinders progress in artistic image synthesis. Existing evaluation metrics, such as Fréchet Inception Distance (FID) and CMMD, struggle with aesthetic assessment: they rely on pretrained visual features that overlook nuanced artistic attributes and employ distance functions ill-suited for modeling the diverse, multi-modal distribution of artistic styles. To address these limitations, we propose ArtFRD, a metric specifically designed for generative aesthetic evaluation. Grounded in aesthetic theory, ArtFRD extracts visual features along four key aesthetic dimensions-brushstroke, composition, lighting, and color-to capture fine-grained artistic properties. To model the multi-modal nature of artistic styles, we adopt a Gaussian Mixture Model assumption and derive an efficient approximation of the Fisher-Rao distance, which serves as the final evaluation score. Extensive experiments demonstrate that ArtFRD aligns significantly better with human aesthetic judgments than existing metrics, even across a wide range of artistic styles. These results highlight its potential as a robust and interpretable foundation for future research in generative aesthetic evaluation.
Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jinchao Zhang 0001, Jie Zhou 0016
ACM Multimedia6
2025 Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
abstract
Multimodal language analysis is a rapidly evolving field that leverages multiple modalities to enhance the understanding of high-level semantics underlying human conversational utterances. Despite its significance, little research has investigated the capability of multimodal large language models (MLLMs) to comprehend cognitive-level semantics. In this paper, we introduce MMLA, a comprehensive benchmark specifically designed to address this gap. MMLA comprises over 61K multimodal utterances drawn from both staged and real-world scenarios, covering six core dimensions of multimodal semantics: intent, emotion, dialogue act, sentiment, speaking style, and communication behavior. We evaluate eight mainstream branches of LLMs and MLLMs using three methods: zero-shot inference, supervised fine-tuning, and instruction tuning. Extensive experiments reveal that even fine-tuned models achieve only about 60~70% accuracy, underscoring the limitations of current MLLMs in understanding complex human language. We believe that MMLA will serve as a solid foundation for exploring the potential of large language models in multimodal language analysis and provide valuable resources to advance this field. The datasets and code are open-sourced at https://github.com/thuiar/MMLA.
Hanlei Zhang, Hua Xu 0003, Yeshuang Zhu, Peiwu Wang, Haige Zhu, Jie Zhou 0016, Jinchao Zhang 0001
NeurIPS8
2024 Enhancing Note-Level Singing Transcription Model with Unlabeled and Weakly Labeled Data
abstract
Note-level automatic singing transcription, involving the extraction of onset, offset, and pitch information from a singing voice, is a crucial process in the field of Music Information Retrieval (MIR), The recent advancements in deep learning models have led to significant progress in this field. However, annotating a training dataset requires professional music expertise, and the entire annotation process is time-consuming and labor-intensive. Therefore, this field suffers from a severe data scarcity problem. To address this issue, we developed a singing transcription model based on wav2vec 2.0, a pretrained speech representation model. The model can learn from unlabeled speech and weakly-labeled singing data and use this knowledge to benefit the transcription task. The experiments showed that our proposed method achieves a significant improvement over previous approaches on various benchmarks. Moreover, additional experiments demonstrate that our method achieves competitive performance even with a small proportion of training data.
Yao Qiu, Jinchao Zhang 0001, Yong Shan, Jie Zhou 0016
ICASSP2
2024 Overview of the Tenth Dialog System Technology Challenge: DSTC10
abstract
This article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks.
Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky
IEEE ACM Trans. Audio Speech Lang. Process.10
2023 Rephrasing the Reference for Non-autoregressive Machine Translation
abstract
Non-autoregressive neural machine translation (NAT) models suffer from the multi-modality problem that there may exist multiple possible translations of a source sentence, so the reference sentence may be inappropriate for the training when the NAT output is closer to other translations. In response to this problem, we introduce a rephraser to provide a better training target for NAT by rephrasing the reference sentence according to the NAT output. As we train NAT based on the rephraser output rather than the reference sentence, the rephraser output should fit well with the NAT output and not deviate too far from the reference, which can be quantified as reward functions and optimized by reinforcement learning. Experiments on major WMT benchmarks and NAT baselines show that our approach consistently improves the translation quality of NAT. Specifically, our best variant achieves comparable performance to the autoregressive Transformer, while being 14.7 times more efficient in inference.
Chenze Shao, Jinchao Zhang 0001, Jie Zhou 0016, Yang Feng 0004
AAAI2
2023 Humming2Music: Being A Composer As Long As You Can Humming
abstract
Creating a piece of music is difficult for people who have never been trained to compose. We present an automatic music generation system to lower the threshold of creating music. The system takes the user's humming as input and creates full music based on the humming melody. The system consists of five modules: 1) humming transcription, 2) melody generation, 3) broken chord generation, 4) accompaniment generation, and 5) audio synthesis. The first module transcribes the user's humming audio to a score, and then the melody generation module composes a complete melody based on the user's humming melody. After that, the third module will generate a broken chord track to accompany the full melody, and the fourth module will create more accompanying tracks. Finally, the audio synthesis module mixes all the tracks to generate the music. Through the user experiment, our system can generate high-quality music with natural expression based on the user's humming input.
Yao Qiu, Jinchao Zhang 0001, Huiying Ren, Yong Shan, Jie Zhou 0016
IJCAI2
2023 LingGe: An Automatic Ancient Chinese Poem-to-Song Generation System
abstract
This paper presents a novel system, named LingGe ("伶歌" in Chinese), to generate songs for ancient Chinese poems automatically. LingGe takes the poem as the lyric, composes music conditioned on the lyric, and finally outputs a full song including the singing and the accompaniment. It consists of four modules: rhythm recognition, melody generation, accompaniment generation, and audio synthesis. Firstly, the rhythm recognition module analyzes the song structure and rhythm according to the poem. Secondly, the melody generation module assembles the rhythm into the template and then generates the melody. Thirdly, the accompaniment generation module predicts the accompaniment in harmony with the melody. Finally, the audio synthesis module generates singing and accompaniment audio and then mixes them to obtain songs. The results show that LingGe can generate high-quality and expressive songs for ancient Chinese poems, both in harmony and rhythm.
Yong Shan, Jinchao Zhang 0001, Huiying Ren, Yao Qiu, Jie Zhou 0016
IJCAI2
2022 Counterfactual Data Augmentation via Perspective Transition for Open-Domain Dialogues
abstract
The construction of open-domain dialogue systems requires high-quality dialogue datasets.The dialogue data admits a wide variety of responses for a given dialogue history, especially responses with different semantics.However, collecting high-quality such a dataset in most scenarios is labor-intensive and timeconsuming.In this paper, we propose a data augmentation method to automatically augment high-quality responses with different semantics by counterfactual inference.Specifically, given an observed dialogue, our counterfactual generation model first infers semantically different responses by replacing the observed reply perspective with substituted ones.Furthermore, our data selection method filters out detrimental augmented responses.Experimental results show that our data augmentation method can augment high-quality responses with different semantics for a given dialogue history, and can outperform competitive baselines on multiple downstream tasks.
Jiao Ou, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016
EMNLP2
2022 High-Resolution and Arbitrary-Sized Chinese Landscape Painting Creation Based on Generative Adversarial Networks
abstract
This paper outlines an automated creation system for Chinese landscape paintings based on generative adversarial networks. The system consists of three cascaded modules: generation, resizing, and super-resolution. The generation module first generates a square-shaped painting, then the resizing module predicts an appropriate aspect ratio for it and performs resizing, and finally the super-resolution module is used to increase the resolution and improve the quality. After training each module with the images we collected from the web, our system can create high-resolution landscape paintings in arbitrary sizes.
Peixiang Luo, Jinchao Zhang 0001, Jie Zhou 0016
IJCAI2
2022 Structure-Enhanced Pop Music Generation via Harmony-Aware Learning
abstract
Pop music generation has always been an attractive topic for both musicians and scientists for a long time. However, automatically composing pop music with a satisfactory structure is still a challenging issue. In this paper, we propose to leverage harmony-aware learning for structure-enhanced pop music generation. On the one hand, one of the participants of harmony, chord, represents the harmonic set of multiple notes, which is integrated closely with the spatial structure of music, the texture. On the other hand, the other participant of harmony, chord progression, usually accompanies the development of the music, which promotes the temporal structure of music, the form. Moreover, when chords evolve into chord progression, the texture and form can be bridged by the harmony naturally, which contributes to the joint learning of the two structures. Furthermore, we propose the Harmony-Aware Hierarchical Music Transformer (HAT), which can exploit the structure adaptively from the music, and make the musical tokens interact hierarchically to enhance the structure in multi-level musical elements. Experimental results reveal that compared to the existing methods, HAT owns a much better understanding of the structure and it can also improve the quality of generated music, especially in the form and texture.
Xueyao Zhang, Jinchao Zhang 0001, Yao Qiu, Jie Zhou 0016
ACM Multimedia2
2021 Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue Utterances
abstract
Zekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zekang Li, Jinchao Zhang 0001, Zhengcong Fei, Yang Feng 0004, Jie Zhou 0016
ACL/IJCNLP (1)2
2021 GTM: A Generative Triple-wise Model for Conversational Question Generation
abstract
Lei Shen, Fandong Meng, Jinchao Zhang, Yang Feng, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Lei Shen 0001, Fandong Meng, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016
ACL/IJCNLP (1)3
2021 Different Strokes for Different Folks: Investigating Appropriate Further Pre-training Approaches for Diverse Dialogue Tasks
abstract
Loading models pre-trained on the large-scale corpus in the general domain and fine-tuning them on specific downstream tasks is gradually becoming a paradigm in Natural Language Processing.Previous investigations prove that introducing a further pre-training phase between pre-training and fine-tuning phases to adapt the model on the domain-specific unlabeled data can bring positive effects.However, most of these further pre-training works just keep running the conventional pre-training task, e.g., masked language model, which can be regarded as the domain adaptation to bridge the data distribution gap.After observing diverse downstream tasks, we suggest that different tasks may also need a further pre-training phase with appropriate training tasks to bridge the task formulation gap.To investigate this, we carry out a study for improving multiple task-oriented dialogue downstream tasks through designing various tasks at the further pre-training phase.The experiment shows that different downstream tasks prefer different further pre-training tasks, which have intrinsic correlation and most further pre-training tasks significantly improve certain target tasks rather than all.Our investigation indicates that it is of great importance and effectiveness to design appropriate further pre-training tasks modeling specific information that benefit downstream tasks.Besides, we present multiple constructive empirical conclusions for enhancing task-oriented dialogues.
Yao Qiu, Jinchao Zhang 0001, Jie Zhou 0016
EMNLP (1)2
2021 Text AutoAugment: Learning Compositional Augmentation Policy for Text Classification
abstract
Data augmentation aims to enrich training samples for alleviating the overfitting issue in low-resource or class-imbalanced situations.Traditional methods first devise task-specific operations such as Synonym Substitute, then preset the corresponding parameters such as the substitution rate artificially, which require a lot of prior knowledge and are prone to fall into the sub-optimum.Besides, the number of editing operations is limited in the previous methods, which decreases the diversity of the augmented data and thus restricts the performance gain.To overcome the above limitations, we propose a framework named Text AutoAugment (TAA) to establish a compositional and learnable paradigm for data augmentation.We regard a combination of various operations as an augmentation policy and utilize an efficient Bayesian Optimization algorithm to automatically search for the best policy, which substantially improves the generalization capability of models.Experiments on six benchmark datasets show that TAA boosts classification accuracy in low-resource and class-imbalanced regimes by an average of 8.8% and 9.7%, respectively, outperforming strong baselines.1
Shuhuai Ren, Jinchao Zhang 0001, Lei Li 0039, Xu Sun 0001, Jie Zhou 0016
EMNLP (1)2
2021 Sequence-Level Training for Non-Autoregressive Neural Machine Translation
abstract
Abstract In recent years, Neural Machine Translation (NMT) has achieved notable results in various translation tasks. However, the word-by-word generation manner determined by the autoregressive mechanism leads to high translation latency of the NMT and restricts its low-latency applications. Non-Autoregressive Neural Machine Translation (NAT) removes the autoregressive mechanism and achieves significant decoding speedup by generating target words independently and simultaneously. Nevertheless, NAT still takes the word-level cross-entropy loss as the training objective, which is not optimal because the output of NAT cannot be properly evaluated due to the multimodality problem. In this article, we propose using sequence-level training objectives to train NAT models, which evaluate the NAT outputs as a whole and correlates well with the real translation quality. First, we propose training NAT models to optimize sequence-level evaluation metrics (e.g., BLEU) based on several novel reinforcement algorithms customized for NAT, which outperform the conventional method by reducing the variance of gradient estimation. Second, we introduce a novel training objective for NAT models, which aims to minimize the Bag-of-N-grams (BoN) difference between the model output and the reference sentence. The BoN training objective is differentiable and can be calculated efficiently without doing any approximations. Finally, we apply a three-stage training strategy to combine these two methods to train the NAT model. We validate our approach on four translation tasks (WMT14 En↔De, WMT16 En↔Ro), which shows that our approach largely outperforms NAT baselines and achieves remarkable performance on all translation tasks. The source code is available at https://github.com/ictnlp/Seq-NAT.
Chenze Shao, Yang Feng 0004, Jinchao Zhang 0001, Fandong Meng, Jie Zhou 0016
Comput. Linguistics3
2021 A dependency syntactic knowledge augmented interactive architecture for end-to-end aspect-based sentiment analysis
Yunlong Liang, Fandong Meng, Jinchao Zhang 0001, Yufeng Chen 0005, Jin An Xu, Jie Zhou 0016
Neurocomputing3
2021 Bridging Text and Video: A Universal Multimodal Transformer for Audio-Visual Scene-Aware Dialog
abstract
Audio-Visual Scene-Aware Dialog (AVSD) is a task to generate responses when chatting about a given video, which is organized as a track of the 8th Dialog System Technology Challenge (DSTC8). There are two challenges in this task: 1) making effective interaction among different modalities; 2) better understanding dialogues and generating informative responses. To tackle the challenges, we propose a universal multimodal transformer and introduce the multi-task learning method to learn joint representations among different modalities as well as generate informative and fluent responses by leveraging the pre-trained language model. Our method extends the natural language generation pre-trained model to multimodal dialogue generation task, which allows fine-tuning language models to capture information across both visual and textual modalities. Our system achieves the best performance in the objective evaluation in both DSTC7-AVSD and DSTC8-AVSD dataset and achieves an impressive 98.4% of the human performance based on human ratings in the DSTC8-AVSD challenge.
Zekang Li, Zongjia Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Multi-Zone Unit for Recurrent Neural Networks
Fandong Meng, Jinchao Zhang 0001, Yang Liu 0005, Jie Zhou 0016
AAAI2
2020 Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine Translation
abstract
Non-Autoregressive Neural Machine Translation (NAT) achieves significant decoding speedup through generating target words independently and simultaneously. However, in the context of non-autoregressive translation, the word-level cross-entropy loss cannot model the target-side sequential dependency properly, leading to its weak correlation with the translation quality. As a result, NAT tends to generate influent translations with over-translation and under-translation errors. In this paper, we propose to train NAT to minimize the Bag-of-Ngrams (BoN) difference between the model output and the reference sentence. The bag-of-ngrams training objective is differentiable and can be efficiently calculated, which encourages NAT to capture the target-side sequential dependency and correlates well with the translation quality. We validate our approach on three translation tasks and show that our approach largely outperforms the NAT baseline by about 5.0 BLEU scores on WMT14 En↔De and about 2.5 BLEU scores on WMT16 En↔Ro.
Chenze Shao, Jinchao Zhang 0001, Yang Feng 0004, Fandong Meng, Jie Zhou 0016
AAAI2
2020 A Contextual Hierarchical Attention Network with Adaptive Objective for Dialogue State Tracking
abstract
Recent studies in dialogue state tracking (DST) leverage historical information to determine states which are generally represented as slot-value pairs. However, most of them have limitations to efficiently exploit relevant context due to the lack of a powerful mechanism for modeling interactions between the slot and the dialogue history. Besides, existing methods usually ignore the slot imbalance problem and treat all slots indiscriminately, which limits the learning of hard slots and eventually hurts overall performance. In this paper, we propose to enhance the DST through employing a contextual hierarchical attention network to not only discern relevant information at both word level and turn level but also learn contextual representations. We further propose an adaptive objective to alleviate the slot imbalance problem by dynamically adjust weights of different slots during training. Experimental results show that our approach reaches 52.68% and 58.55% joint accuracy on MultiWOZ 2.0 and MultiWOZ 2.1 datasets respectively and achieves new state-of-the-art performance with considerable improvements (+1.24% and +5.98%).
Yong Shan, Zekang Li, Jinchao Zhang 0001, Fandong Meng, Yang Feng 0004, Cheng Niu, Jie Zhou 0016
ACL3
2020 Contrastive Zero-Shot Learning for Cross-Domain Slot Filling with Adversarial Attack
abstract
Zero-shot slot filling has widely arisen to cope with data scarcity in target domains.However, previous approaches often ignore constraints between slot value representation and related slot description representation in the latent space and lack enough model robustness.In this paper, we propose a Contrastive Zero-Shot Learning with Adversarial Attack (CZSL-Adv) method for the cross-domain slot filling.The contrastive loss aims to map slot value contextual representations to the corresponding slot description representations.And we introduce an adversarial attack training strategy to improve model robustness.Experimental results show that our model significantly outperforms state-of-the-art baselines under both zero-shot and few-shot settings.
Keqing He 0001, Jinchao Zhang 0001, Yuanmeng Yan, Weiran Xu, Cheng Niu, Jie Zhou 0016
COLING2
2020 One Comment from One Perspective: An Effective Strategy for Enhancing Automatic Music Comment
abstract
The automatic generation of music comments is of great significance for increasing the popularity of music and the music platform's activity.In human music comments, there exists high distinction and diverse perspectives for the same song.In other words, for a song, different comments stem from different musical perspectives.However, to date, this characteristic has not been considered well in research on automatic comment generation.The existing methods tend to generate common and meaningless comments.In this paper, we propose an effective multiperspective strategy to enhance the diversity of the generated comments.The experiment results on two music comment datasets show that our proposed model can effectively generate a series of diverse music comments based on different perspectives, which outperforms state-of-the-art baselines by a substantial margin. 1
Tengfei Huo, Jinchao Zhang 0001, Jie Zhou 0016
COLING3
2020 Token-level Adaptive Training for Neural Machine Translation
abstract
There exists a token imbalance phenomenon in natural language as different tokens appear with different frequencies, which leads to different learning difficulties for tokens in Neural Machine Translation (NMT).The vanilla NMT model usually adopts trivial equal-weighted objectives for target tokens with different frequencies and tends to generate more high-frequency tokens and less lowfrequency tokens compared with the golden token distribution.However, low-frequency tokens may carry critical semantic information that will affect the translation quality once they are neglected.In this paper, we explored target token-level adaptive objectives based on token frequencies to assign appropriate weights for each target token during training.We aimed that those meaningful but relatively low-frequency words could be assigned with larger weights in objectives to encourage the model to pay more attention to these tokens.Our method yields consistent improvements in translation quality on ZH-EN, EN-RO, and EN-DE translation tasks, especially on sentences that contain more low-frequency tokens where we can get 1.68, 1.02, and 0.52 BLEU increases compared with baseline, respectively.Further analyses show that our method can also improve the lexical diversity of translation.
Shuhao Gu, Jinchao Zhang 0001, Fandong Meng, Yang Feng 0004, Wanying Xie, Jie Zhou 0016, Dong Yu 0003
EMNLP (1)2
2019 DTMT: A Novel Deep Transition Architecture for Neural Machine Translation
abstract
Past years have witnessed rapid developments in Neural Machine Translation (NMT). Most recently, with advanced modeling and training techniques, the RNN-based NMT (RNMT) has shown its potential strength, even compared with the well-known Transformer (self-attentional) model. Although the RNMT model can possess very deep architectures through stacking layers, the transition depth between consecutive hidden states along the sequential axis is still shallow. In this paper, we further enhance the RNN-based NMT through increasing the transition depth between consecutive hidden states and build a novel Deep Transition RNN-based Architecture for Neural Machine Translation, named DTMT. This model enhances the hidden-to-hidden transition with multiple non-linear transformations, as well as maintains a linear transformation path throughout this deep transition by the well-designed linear transformation mechanism to alleviate the gradient vanishing problem. Experiments show that with the specially designed deep transition modules, our DTMT can achieve remarkable improvements on translation quality. Experimental results on Chinese⇒English translation task show that DTMT can outperform the Transformer model by +2.09 BLEU points and achieve the best results ever reported in the same dataset. On WMT14 English⇒German and English⇒French translation tasks, DTMT shows superior quality to the state-of-the-art NMT systems, including the Transformer and the RNMT+.
Fandong Meng, Jinchao Zhang 0001
AAAI2
2019 GCDT: A Global Context Enhanced Deep Transition Architecture for Sequence Labeling
abstract
Current state-of-the-art systems for the sequence labeling tasks are typically based on the family of Recurrent Neural Networks (RNNs).However, the shallow connections between consecutive hidden states of RNNs and insufficient modeling of global information restrict the potential performance of those models.In this paper, we try to address these issues, and thus propose a Global Context enhanced Deep Transition architecture for sequence labeling named GCDT.We deepen the state transition path at each position in a sentence, and further assign every token with a global representation learned from the entire sentence.Experiments on two standard sequence labeling tasks show that, given only training data and the ubiquitous word embeddings (Glove), our GCDT achieves 91.96 F 1 on the CoNLL03 NER task and 95.43 F 1 on the CoNLL2000 Chunking task, which outperforms the best reported results under the same settings.Furthermore, by leveraging BERT as an additional resource, we establish new stateof-the-art results with 93.47 F 1 on NER and 97.30F 1 on Chunking 1 .
Yijin Liu, Fandong Meng, Jinchao Zhang 0001, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
ACL (1)3
2019 Retrieving Sequential Information for Non-Autoregressive Neural Machine Translation
abstract
Non-Autoregressive Transformer (NAT) aims to accelerate the Transformer model through discarding the autoregressive mechanism and generating target words independently, which fails to exploit the target sequential information.Over-translation and under-translation errors often occur for the above reason, especially in the long sentence translation scenario.In this paper, we propose two approaches to retrieve the target sequential information for NAT to enhance its translation ability while preserving the fast-decoding property.Firstly, we propose a sequence-level training method based on a novel reinforcement algorithm for NAT (Reinforce-NAT) to reduce the variance and stabilize the training procedure.Secondly, we propose an innovative Transformer decoder named FS-decoder to fuse the target sequential information into the top layer of the decoder.Experimental results on three translation tasks show that the Reinforce-NAT surpasses the baseline NAT system by a significant margin on BLEU without decelerating the decoding speed and the FS-decoder achieves comparable translation performance to the autoregressive Transformer with considerable speedup.
Chenze Shao, Yang Feng 0004, Jinchao Zhang 0001, Fandong Meng, Xilin Chen 0001, Jie Zhou 0016
ACL (1)3
2019 A Novel Aspect-Guided Deep Transition Model for Aspect Based Sentiment Analysis
abstract
Yunlong Liang, Fandong Meng, Jinchao Zhang, Jinan Xu, Yufeng Chen, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yunlong Liang, Fandong Meng, Jinchao Zhang 0001, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016
EMNLP/IJCNLP (1)3
2019 CM-Net: A Novel Collaborative Memory Network for Spoken Language Understanding
abstract
Yijin Liu, Fandong Meng, Jinchao Zhang, Jie Zhou, Yufeng Chen, Jinan Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yijin Liu, Fandong Meng, Jinchao Zhang 0001, Jie Zhou 0016, Yufeng Chen 0005, Jin An Xu
EMNLP/IJCNLP (1)3
2019 Enhancing Context Modeling with a Query-Guided Capsule Network for Document-level Translation
abstract
Zhengxin Yang, Jinchao Zhang, Fandong Meng, Shuhao Gu, Yang Feng, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zhengxin Yang, Jinchao Zhang 0001, Fandong Meng, Shuhao Gu, Yang Feng 0004, Jie Zhou 0016
EMNLP/IJCNLP (1)2
2017 Incorporating Word Reordering Knowledge into Attention-based Neural Machine Translation
abstract
This paper proposes three distortion models to explicitly incorporate the word reordering knowledge into attention-based Neural Machine Translation (NMT) for further improving translation performance.Our proposed models enable attention mechanism to attend to source words regarding both the semantic requirement and the word reordering penalty.Experiments on Chinese-English translation show that the approaches can improve word alignment quality and achieve significant translation improvements over a basic attention-based N-MT by large margins.Compared with previous works on identical corpora, our system achieves the state-of-the-art performance on translation quality.
Jinchao Zhang 0001, Mingxuan Wang, Qun Liu 0001, Jie Zhou 0016
ACL (1)1
2017 ME-MD: An Effective Framework for Neural Machine Translation with Multiple Encoders and Decoders
abstract
The encoder-decoder neural framework is widely employed for Neural Machine Translation (NMT) with a single encoder to represent the source sentence and a single decoder to generate target words. The translation performance heavily relies on the representation ability of the encoder and the generation ability of the decoder. To further enhance NMT, we propose to extend the original encoder-decoder framework to a novel one, which has multiple encoders and decoders (ME-MD). Through this way, multiple encoders extract more diverse features to represent the source sequence and multiple decoders capture more complicated translation knowledge. Our proposed ME-MD framework is convenient to integrate heterogeneous encoders and decoders with multiple depths and multiple types. Experiment on Chinese-English translation task shows that our ME-MD system surpasses the state-of-the-art NMT system by 2.1 BLEU points and surpasses the phrase-based Moses by 7.38 BLEU points. Our framework is general and can be applied to other sequence to sequence tasks.
Jinchao Zhang 0001, Qun Liu 0001, Jie Zhou 0016
IJCAI1
2014 BigOP: Generating Comprehensive Big Data Workloads as a Benchmarking Framework
Yuqing Zhu 0001, Jianfeng Zhan, Chuliang Weng, Raghunath Othayoth Nambiar, Jinchao Zhang 0001, Xingzhen Chen, Lei Wang 0004
DASFAA (2)5