EDBT 2026 Demo / reviewers in the wild / expert
Zhaojiang Lin
dblp:228/9217
· DBLP profile ↗
27ranked-venue papers
7as first author
14since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Extrapolating to Unknown Opinions Using LLMsabstractFrom ice cream flavors to climate change, people exhibit a wide array of opinions on various topics, and understanding the rationale for these opinions can promote healthy discussion and consensus among them. As such, it can be valuable for a large language model (LLM), particularly as an AI assistant, to be able to empathize with or even explain these various standpoints. In this work, we hypothesize that different topic stances often manifest correlations that can be used to extrapolate to topics with unknown opinions. We explore various prompting and fine-tuning methods to improve an LLM’s ability to (a) extrapolate from opinions on known topics to unknown ones and (b) support their extrapolation with reasoning. Our findings suggest that LLMs possess inherent knowledge from training data about these opinion correlations, and with minimal data, the similarities between human opinions and model-extrapolated opinions can be improved by more than 50%. Furthermore, LLM can generate the reasoning process behind their extrapolation of opinions. Kexun Zhang, Jane Dwivedi-Yu, Zhaojiang Lin, Yuning Mao, William Yang Wang, Lei Li 0005, Yi-Chia Wang |
COLING | 3 |
| 2025 | Proactive Assistant Dialogue Generation from Streaming Egocentric VideosabstractYichi Zhang, Xin Luna Dong, Zhaojiang Lin, Andrea Madotto, Anuj Kumar, Babak Damavandi, Joyce Chai, Seungwhan Moon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yichi Zhang 0001, Xin Dong 0001, Zhaojiang Lin, Andrea Madotto, Babak Damavandi, Joyce Y. Chai, Seungwhan Moon |
EMNLP | 3 |
| 2025 | Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
Jiamin Xie, Ju Lin, Yiteng Huang, Tyler Vuong, Zhaojiang Lin, Prashant Rawat, Sangeeta Srivastava, Ming Sun 0013, Florian Metze |
INTERSPEECH | 5 |
| 2025 | VisualLens: Personalization through Task-Agnostic Visual HistoryabstractExisting recommendation systems either rely on user interaction logs, such as online shopping history for shopping recommendations, or focus on text signals.
However, item-based histories are not always accessible and generalizable for multimodal recommendation.
We hypothesize that a user's visual history --- comprising images from daily life --- can offer rich, task-agnostic insights into their interests and preferences, and thus be leveraged for effective personalization.
To this end, we propose VisualLens, a novel framework that leverages multimodal large language models (MLLMs) to enable personalization using task-agnostic visual history.
VisualLens extracts, filters, and refines a spectrum user profile from the visual history to support personalized recommendation.
We created two new benchmarks, Google-Review-V and Yelp-V, with task-agnostic visual histories, and show that VisualLens improves over state-of-the-art item-based multimodal recommendations by 5-10\% on Hit@3, and outperforms GPT-4o by 2-5\%.
Further analysis shows that VisualLens is robust across varying history lengths and excels at adapting to both longer histories and unseen content categories. Wang Zhu 0001, Deqing Fu, Kai Sun 0006, Zhaojiang Lin, Seungwhan Moon, Kanika Narang, Mustafa Canim, Xin Dong 0001 |
NeurIPS | 5 |
| 2025 | Corgi: Cached Memory Guided Video GenerationabstractText-to-Video generation has achieved remarkable progress with the rise of diffusion models. In this work, we introduce Cached Memory-Guided Video Generation (Corgi), aiming to generate multi-scene videos with arbi-trary number of video clips, conditioned on input images and instruction prompts. This is a challenging task, as tra-ditional T2V methods often struggle to maintain the quality of longer videos due to the difficulties in preserving visual context from earlier scenes. We address this by introducing a cached memory mechanism that stores the key frames. Our multi-scene video generation process is explicitly con-ditioned on the cached memories to avoid forgetting the vi-sual appearance of target subjects. Corgi shows significant improvement in multi-scene video generation compared to the prior art, with up to 59.2% in long-term consistency and 7.6% in diversity. Xindi Wu, Uriel Singer, Zhaojiang Lin, Andrea Madotto, Xide Xia, Paul A. Crook, Xin Dong 0001, Seungwhan Moon |
WACV | 3 |
| 2024 | Large Language Models as Zero-shot Dialogue State Tracker through Function CallingabstractZekun Li, Zhiyu Zoey Chen, Mike Ross, Patrick Huber, Seungwhan Moon, Zhaojiang Lin, Luna Dong, Adithya Sagar, Xifeng Yan, Paul A. Crook. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zekun Li 0001, Zhiyu Chen 0002, Mike Ross, Patrick Huber, Seungwhan Moon, Zhaojiang Lin, Xin Dong 0001, Adithya Sagar, Xifeng Yan, Paul A. Crook |
ACL (1) | 6 |
| 2023 | Introducing Semantics into Speech EncodersabstractDerek Xu, Shuyan Dong, Changhan Wang, Suyoun Kim, Zhaojiang Lin, Bing Liu, Akshat Shrivastava, Shang-Wen Li, Liang-Hsuan Tseng, Guan-Ting Lin, Alexei Baevski, Hung-yi Lee, Yizhou Sun, Wei Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Derek Xu, Shuyan Dong, Changhan Wang, Suyoun Kim, Zhaojiang Lin, Akshat Shrivastava, Shang-Wen Li 0001, Liang-Hsuan Tseng, Guan-Ting Lin, Alexei Baevski, Hung-yi Lee, Yizhou Sun, Wei Wang 0010 |
ACL (1) | 5 |
| 2023 | Continual Dialogue State Tracking via Example-Guided Question AnsweringabstractHyundong Cho, Andrea Madotto, Zhaojiang Lin, Khyathi Chandu, Satwik Kottur, Jing Xu, Jonathan May, Chinnadhurai Sankar. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Hyundong Cho, Andrea Madotto, Zhaojiang Lin, Khyathi Raghavi Chandu, Satwik Kottur, Jonathan May, Chinnadhurai Sankar |
EMNLP | 3 |
| 2022 | FaceFormer: Speech-Driven 3D Facial Animation with TransformersabstractSpeech-driven 3D facial animation is challenging due to the complex geometry of human faces and the limited availability of 3D audio-visual data. Prior works typically focus on learning phoneme-level features of short audio windows with limited context, occasionally resulting in inaccurate lip movements. To tackle this limitation, we propose a Transformer-based autoregressive model, Face-Former, which encodes the long-term audio context and autoregressively predicts a sequence of animated 3D face meshes. To cope with the data scarcity issue, we integrate the self-supervised pre-trained speech representations. Also, we devise two biased attention mechanisms well suited to this specific task, including the biased cross-modal multi-head (MH) attention and the biased causal MH self-attention with a periodic positional encoding strategy. The former effectively aligns the audio-motion modalities, whereas the latter offers abilities to generalize to longer audio sequences. Extensive experiments and a perceptual user study show that our approach outperforms the existing state-of-the-arts. The code and the video are available at: https://evelynfan.github.io/audio2face/ Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang 0001, Taku Komura |
CVPR | 2 |
| 2021 | The Adapter-Bot: All-In-One Controllable Conversational ModelabstractIn this paper, we present the Adapter-Bot, a generative chat-bot that uses a fixed backbone conversational model such as DialGPT (Zhang et al. 2019) and triggers on-demand dialogue skills via different adapters (Houlsby et al. 2019). Each adapter can be trained independently, thus allowing a continual integration of skills without retraining the entire model. Depending on the skills, the model is able to process multiple knowledge types, such as text, tables, and graphs, in a seamless manner. The dialogue skills can be triggered automatically via a dialogue manager, or manually, thus allowing high-level control of the generated responses. At the current stage, we have implemented 12 response styles (e.g., positive, negative etc.), 6 goal-oriented skills (e.g. weather information, movie recommendation, etc.), and personalized and emphatic responses. Zhaojiang Lin, Andrea Madotto, Yejin Bang, Pascale Fung |
AAAI | 1 |
| 2021 | On the Importance of Word Order Information in Cross-lingual Sequence LabelingabstractCross-lingual models trained on source language tasks possess the capability to directly transfer to target languages. However, since word order variances generally exist in different languages, cross-lingual models that overfit into the word order of the source language could have sub-optimal performance in target languages. In this paper, we hypothesize that reducing the word order information fitted into the models can improve the adaptation performance in target languages. To verify this hypothesis, we introduce several methods to make models encode less word order information of the source language and test them based on cross-lingual word embeddings and the pre-trained multilingual model. Experimental results on three sequence labeling tasks (i.e., part-of-speech tagging, named entity recognition and slot filling tasks) show that reducing word order information injected into the model can achieve better zero-shot cross-lingual performance. Further analysis illustrates that fitting excessive or insufficient word order information into the model results in inferior cross-lingual performance. Moreover, our proposed methods can also be applied to strong cross-lingual models and further improve their performance. Zihan Liu 0001, Genta Indra Winata, Samuel Cahyawijaya, Andrea Madotto, Zhaojiang Lin, Pascale Fung |
AAAI | 5 |
| 2021 | Zero-Shot Dialogue State Tracking via Cross-Task TransferabstractZhaojiang Lin, Bing Liu, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul Crook, Zhiguang Wang, Zhou Yu, Eunjoon Cho, Rajen Subba, Pascale Fung. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhaojiang Lin, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul A. Crook, Zhiguang Wang, Zhou Yu 0005, Eunjoon Cho, Rajen Subba, Pascale Fung |
EMNLP (1) | 1 |
| 2021 | Continual Learning in Task-Oriented Dialogue SystemsabstractAndrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul Crook, Bing Liu, Zhou Yu, Eunjoon Cho, Pascale Fung, Zhiguang Wang. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Andrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul A. Crook, Zhou Yu 0005, Eunjoon Cho, Pascale Fung, Zhiguang Wang |
EMNLP (1) | 2 |
| 2021 | Leveraging Slot Descriptions for Zero-Shot Cross-Domain Dialogue StateTrackingabstractZhaojiang Lin, Bing Liu, Seungwhan Moon, Paul Crook, Zhenpeng Zhou, Zhiguang Wang, Zhou Yu, Andrea Madotto, Eunjoon Cho, Rajen Subba. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhaojiang Lin, Seungwhan Moon, Paul A. Crook, Zhenpeng Zhou, Zhiguang Wang, Andrea Madotto, Eunjoon Cho, Rajen Subba |
NAACL-HLT | 1 |
| 2020 | CAiRE: An End-to-End Empathetic ChatbotabstractWe present CAiRE, an end-to-end generative empathetic chatbot designed to recognize user emotions and respond in an empathetic manner. Our system adapts the Generative Pre-trained Transformer (GPT) to empathetic response generation task via transfer learning. CAiRE is built primarily to focus on empathy integration in fully data-driven generative dialogue systems. We create a web-based user interface which allows multiple users to asynchronously chat with CAiRE. CAiRE also collects user feedback and continues to improve its response quality by discarding undesirable generations via active learning and negative training. Zhaojiang Lin, Peng Xu 0008, Genta Indra Winata, Farhad Bin Siddique, Zihan Liu 0001, Jamin Shin, Pascale Fung |
AAAI | 1 |
| 2020 | Attention-Informed Mixed-Language Training for Zero-Shot Cross-Lingual Task-Oriented Dialogue SystemsabstractRecently, data-driven task-oriented dialogue systems have achieved promising performance in English. However, developing dialogue systems that support low-resource languages remains a long-standing challenge due to the absence of high-quality data. In order to circumvent the expensive and time-consuming data collection, we introduce Attention-Informed Mixed-Language Training (MLT), a novel zero-shot adaptation method for cross-lingual task-oriented dialogue systems. It leverages very few task-related parallel word pairs to generate code-switching sentences for learning the inter-lingual semantics across languages. Instead of manually selecting the word pairs, we propose to extract source words based on the scores computed by the attention layer of a trained English task-related model and then generate word pairs using existing bilingual dictionaries. Furthermore, intensive experiments with different cross-lingual embeddings demonstrate the effectiveness of our approach. Finally, with very few word pairs, our model achieves significant zero-shot adaptation performance improvements in both cross-lingual dialogue state tracking and natural language understanding (i.e., intent detection and slot filling) tasks compared to the current state-of-the-art approaches, which utilize a much larger amount of bilingual data. Zihan Liu 0001, Genta Indra Winata, Zhaojiang Lin, Peng Xu 0008, Pascale Fung |
AAAI | 3 |
| 2020 | Meta-Transfer Learning for Code-Switched Speech RecognitionabstractAn increasing number of people in the world today speak a mixed-language as a result of being multilingual.However, building a speech recognition system for code-switching remains difficult due to the availability of limited resources and the expense and significant effort required to collect mixed-language data.We therefore propose a new learning method, meta-transfer learning, to transfer learn on a code-switched speech recognition system in a low-resource setting by judiciously extracting information from high-resource monolingual datasets.Our model learns to recognize individual languages, and transfer them so as to better recognize mixed-language speech by conditioning the optimization on the codeswitching data.Based on experimental results, our model outperforms existing baselines on speech recognition and language modeling tasks, and is faster to converge. Genta Indra Winata, Samuel Cahyawijaya, Zhaojiang Lin, Zihan Liu 0001, Peng Xu 0008, Pascale Fung |
ACL | 3 |
| 2020 | MinTL: Minimalist Transfer Learning for Task-Oriented Dialogue SystemsabstractIn this paper, we propose Minimalist Transfer Learning (MinTL) to simplify the system design process of task-oriented dialogue systems and alleviate the over-dependency on annotated data.MinTL is a simple yet effective transfer learning framework, which allows us to plug-and-play pre-trained seq2seq models, and jointly learn dialogue state tracking and dialogue response generation.Unlike previous approaches, which use a copy mechanism to "carryover" the old dialogue states to the new one, we introduce Levenshtein belief spans (Lev), that allows efficient dialogue state tracking with a minimal generation length.We instantiate our learning framework with two pretrained backbones: T5 (Raffel et al., 2019) and BART (Lewis et al., 2019), and evaluate them on MultiWOZ.Extensive experiments demonstrate that: 1) our systems establish new state-of-the-art results on end-to-end response generation, 2) MinTL-based systems are more robust than baseline methods in the low resource setting, and they achieve competitive results with only 20% training data, and 3) Lev greatly improves the inference efficiency 1 . Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, Pascale Fung |
EMNLP (1) | 1 |
| 2020 | Cross-lingual Spoken Language Understanding with Regularized Representation AlignmentabstractDespite the promising results of current crosslingual models for spoken language understanding systems, they still suffer from imperfect cross-lingual representation alignments between the source and target languages, which makes the performance sub-optimal.To cope with this issue, we propose a regularization approach to further align word-level and sentence-level representations across languages without any external resource.First, we regularize the representation of user utterances based on their corresponding labels.Second, we regularize the latent variable model (Liu et al., 2019a) by leveraging adversarial training to disentangle the latent variables.Experiments on the cross-lingual spoken language understanding task show that our model outperforms current state-of-the-art methods in both few-shot and zero-shot scenarios, and our model, trained on a few-shot setting with only 3% of the target language training data, achieves comparable performance to the supervised training with all the training data. 1 Zihan Liu 0001, Genta Indra Winata, Peng Xu 0008, Zhaojiang Lin, Pascale Fung |
EMNLP (1) | 4 |
| 2020 | Lightweight and Efficient End-To-End Speech Recognition Using Low-Rank TransformerabstractHighly performing deep neural networks come at the cost of computational complexity that limits their practicality for deployment on portable devices. We propose the low-rank transformer (LRT), a memory-efficient and fast neural architecture that significantly reduces the parameters and boosts the speed of training and inference for end-to-end speech recognition. Our approach reduces the number of parameters of the network by more than 50% and speeds up the inference time by around 1.35x compared to the baseline transformer model. The experiments show that our LRT model generalizes better and yields lower error rates on both validation and test sets compared to an uncompressed transformer model. The LRT model outperforms those from existing works on several datasets in an end-to-end setting without using an external language model or acoustic data. Genta Indra Winata, Samuel Cahyawijaya, Zhaojiang Lin, Zihan Liu 0001, Pascale Fung |
ICASSP | 3 |
| 2020 | G2RL: Geometry-Guided Representation Learning for Facial Action Unit Intensity EstimationabstractFacial action unit (AU) intensity estimation aims to measure the intensity of different facial muscle movements. The external knowledge such as AU co-occurrence relationship is typically leveraged to improve performance. However, the AU characteristics may vary among individuals due to different physiological structures of human faces. To this end, we propose a novel geometry-guided representation learning (G2RL) method for facial AU intensity estimation. Specifically, our backbone model is based on a heatmap regression framework, where the produced heatmaps reflect rich information associated with AU intensities and their spatial distributions. Besides, we incorporate the external geometric knowledge into the backbone model to guide the training process via a learned projection matrix. The experimental results on two benchmark datasets demonstrate that our method is comparable with the state-of-the-art approaches, and validate the effectiveness of incorporating external geometric knowledge for facial AU intensity estimation. Yingruo Fan, Zhaojiang Lin |
IJCAI | 2 |
| 2020 | Learning Fast Adaptation on Cross-Accented Speech RecognitionabstractLocal dialects influence people to pronounce words of the same language differently from each other. The great variability and complex characteristics of accents create a major challenge for training a robust and accent-agnostic automatic speech recognition (ASR) system. In this paper, we introduce a cross-accented English speech recognition task as a benchmark for measuring the ability of the model to adapt to unseen accents using the existing CommonVoice corpus. We also propose an accent-agnostic approach that extends the model-agnostic meta-learning (MAML) algorithm for fast adaptation to unseen accents. Our approach significantly outperforms joint training in both zero-shot, few-shot, and all-shot in the mixed-region and cross-region settings in terms of word error rate. Copyright © 2020 ISCA Genta Indra Winata, Samuel Cahyawijaya, Zihan Liu 0001, Zhaojiang Lin, Andrea Madotto, Peng Xu 0008, Pascale Fung |
INTERSPEECH | 4 |
| 2020 | Getting To Know You: User Attribute Extraction from DialoguesabstractUser attributes provide rich and useful information for user understanding, yet structured and easy-to-use attributes are often sparsely populated. In this paper, we leverage dialogues with conversational agents, which contain strong suggestions of user information, to automatically extract user attributes. Since no existing dataset is available for this purpose, we apply distant supervision to train our proposed two-stage attribute extractor, which surpasses several retrieval and generation baselines on human evaluation. Meanwhile, we discuss potential applications (e.g., personalized recommendation and dialogue systems) of such extracted user attributes, and point out current limitations to cast light on future work. Chien-Sheng Wu, Andrea Madotto, Zhaojiang Lin, Peng Xu 0008, Pascale Fung |
LREC | 3 |
| 2019 | Personalizing Dialogue Agents via Meta-LearningabstractExisting personalized dialogue models use human designed persona descriptions to improve dialogue consistency.Collecting such descriptions from existing dialogues is expensive and requires hand-crafted feature designs.In this paper, we propose to extend Model-Agnostic Meta-Learning (MAML) (Finn et al., 2017) to personalized dialogue learning without using any persona descriptions.Our model learns to quickly adapt to new personas by leveraging only a few dialogue samples collected from the same user, which is fundamentally different from conditioning the response on the persona descriptions.Empirical results on Persona-chat dataset (Zhang et al., 2018) indicate that our solution outperforms non-metalearning baselines using automatic evaluation metrics, and in terms of human-evaluated fluency and consistency. Andrea Madotto, Zhaojiang Lin, Chien-Sheng Wu, Pascale Fung |
ACL (1) | 2 |
| 2019 | MoEL: Mixture of Empathetic ListenersabstractZhaojiang Lin, Andrea Madotto, Jamin Shin, Peng Xu, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhaojiang Lin, Andrea Madotto, Jamin Shin, Peng Xu 0008, Pascale Fung |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Hierarchical Meta-Embeddings for Code-Switching Named Entity RecognitionabstractGenta Indra Winata, Zhaojiang Lin, Jamin Shin, Zihan Liu, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Genta Indra Winata, Zhaojiang Lin, Jamin Shin, Zihan Liu 0001, Pascale Fung |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Learning Comment Generation by Leveraging User-generated DataabstractExisting models on open-domain comment generation are difficult to train, and they produce repetitive and uninteresting responses. The problem is due to multiple and contradictory responses from a single article, and by the rigidity of retrieval methods. To solve this problem, we propose a combined approach to retrieval and generation methods. We propose an attentive scorer to retrieve informative and relevant comments by leveraging user-generated data. Then, we use such comments, together with the article, as input for a sequence-to-sequence model with copy mechanism. We show the robustness of our model and how it can alleviate the aforementioned issue by using a large scale comment generation dataset. The result shows that the proposed generative model significantly outperforms strong baseline such as Seq2Seq with attention and Information Retrieval models by around 27 and 30 BLEU-1 points respectively. Zhaojiang Lin, Genta Indra Winata, Pascale Fung |
ICASSP | 1 |