Zeming Liu

dblp:256/2166 · DBLP profile ↗
← Back
37ranked-venue papers
8as first author
35since 2021 · last 2026
0000-0002-3691-8097ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 4 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Mem²Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
abstract
Zihao Cheng, Zeming Liu, Yingyu Shan, Xinyi Wang, Xiangrong Zhu, Yunpu Ma, Hongru Wang, Yuhang Guo, Wei Lin, Yunhong Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zeming Liu, Yingyu Shan, Xiangrong Zhu 0002, Yunpu Ma, Hongru Wang 0003, Yuhang Guo 0001, Yunhong Wang 0001
ACL (1)2
2026 PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception
abstract
Embodied Action Sequence Planning focuses on the capability of embodied agents to implement action planning via environmental perception.This technology enables diverse intelligent assistance for real-world scenarios such as home and office environments.To address the limitations of existing embodied agents in meeting the requirement for proactivity and achieving joint understanding of visual and audio information, this study investigates the ability of embodied agents to proactively provide assistance through action sequence planning based on joint understanding of vision and audio perception without explicit human instructions.Correspondingly, we propose PEAP, the first multimodal proactive embodied action sequence planning dataset.We evaluate the performance of multiple Large Language Models on the PEAP dataset.The results demonstrate that these models still exhibit significant deficiencies on this task particularly lacking accurate environmental perception capabilities.Furthermore, ablation experiment and replacement experiment further corroborate that the joint understanding of multimodal information can significantly improve the models' performance on proactive embodied action sequence planning task.Our dataset and code are publicly available 1 .
Tianwei Lan, Zeming Liu, Zhaoxin Fan, Haifeng Wang 0001, Yuhang Guo 0001
ACL (1)3
2026 Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
abstract
Yanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang, Wenbo Yu, Dawei Song, Jie Tang, Yuhang Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang, Jie Tang 0001, Yuhang Guo 0001
ACL (1)3
2026 A3Bench: an audience-aligned multilingual benchmark for video audience insights understanding
Yiming Lei 0001, Guozhen Peng, Zeming Liu, Hui Qiu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu 0001, Annan Li, Yunhong Wang 0001
Frontiers Comput. Sci.3
2026 Towards multi-language repository-level code generation: From-scratch to guided tasks
Silin Li, Zeming Liu, Yuhang Guo 0001, Yuanfang Guo, Yunhong Wang 0001, Haifeng Wang 0001
Neurocomputing3
2026 Characterization of the heterogeneity in SARS-CoV-2 fitness dynamics via graph representation learning
abstract
Understanding the heterogeneity of population-level viral fitness dynamics, which reflect the interplay between intrinsic viral properties and population immunity, is critical for pandemic preparedness. However, how these dynamics vary across diverse immune backgrounds and mutational landscapes remain poorly characterized. We present Geno-GNN, a graph representation learning approach for retrospectively characterizing the viral fitness dynamics of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Geno-GNN accurately predicts angiotensin-converting enzyme 2 (ACE2) binding affinity and immune escape potential across multiple external datasets. Using Geno-GNN, we identified temporal patterns in SARS-CoV-2 fitness and detected varying rates of fitness change associated with distinct immune backgrounds. Virtual mutation scanning revealed two fitness trajectories: broad immune evasion at the cost of ACE2 affinity and ACE2 affinity maintenance at or above the Wuhan-Hu-1 level along with moderate immune escape. Notably, real-world SARS-CoV-2 variants predominantly followed the latter trajectory, sustaining ACE2 affinity via fixed mutations. These findings underscore the heterogeneous, immune-contextualized nature of viral fitness dynamics and the complex evolutionary pathways of SARS-CoV-2.
Zengmiao Wang, Ziqin Zhou, Junfu Wang, Lingyue Yang, Zhirui Zhang, Weina Xu, Zeming Liu, Yuxi Ge, Liang Yang 0002, Quanyi Wang, Yunlong Cao, Yuanfang Guo, Huaiyu Tian
PLoS Comput. Biol.7
2026 Learn more, forget less: A gradient-Aware data selection approach for LLM
Zeming Liu, Yibai Liu, Zheming Song, Qingjie Liu 0001, Guangxu Chen, Yunhong Wang 0001
Signal Process.1
2025 Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt Learning
abstract
With the malicious use and dissemination of multi-modal deepfake videos, researchers start to investigate multi-modal deepfake detection. Unfortunately, most of the existing methods tune all the parameters of the deep network with limited speech video datasets and are trained under coarse-grained consistency supervision, which hinders their generalization ability in practical scenarios. To solve these problems, in this paper, we propose the first multi-task audio-visual prompt learning method for multi-modal deepfake video detection, by exploiting multiple foundation models. Specifically, we construct a two-stream multi-task learning architecture and propose sequential visual prompts and short-time audio prompts to extract multi-modal features, which are aligned at the frame level and utilized in subsequent fine-grained feature matching and fusion. Due to the natural alignment of visual content and audio signal in real data, we propose a frame-level cross-modal feature matching loss function to learn the fine-grained audio-visual consistency. Comprehensive experiments demonstrate the effectiveness and superior generalization ability of our method against the state-of-the-art methods.
Yuanfang Guo, Zeming Liu, Yunhong Wang 0001
AAAI3
2025 STAMPsy: Towards SpatioTemporal-Aware Mixed-Type Dialogues for Psychological Counseling
abstract
Online psychological counseling dialogue systems are trending, offering a convenient and accessible alternative to traditional in-person therapy. However, existing psychological counseling dialogue systems mainly focus on basic empathetic dialogue or QA with minimal professional knowledge and without goal guidance. In many real-world counseling scenarios, clients often seek multi-type help, such as diagnosis, consultation, therapy, console, and common questions, but existing dialogue systems struggle to combine different dialogue types naturally. In this paper, we identify this challenge as how to construct mixed-type dialogue systems for psychological counseling that enable clients to clarify their goals before proceeding with counseling. To mitigate the challenge, we collect a mixed-type counseling dialogues corpus termed STAMPsy, covering five dialogue types, task-oriented dialogue for diagnosis, knowledge-grounded dialogue, conversational recommendation, empathetic dialogue, and question answering, over 5,000 conversations. Moreover, spatiotemporal-aware knowledge enables systems to have world awareness and has been proven to affect one's mental health. Therefore, we link dialogues in STAMPsy to spatiotemporal state and propose a spatiotemporal-aware mixed-type psychological counseling dataset. Additionally, we build baselines on STAMPsy and develop an iterative self-feedback psychological dialogue generation framework, named Self-STAMPsy. Results indicate that clarifying dialogue goals in advance and utilizing spatiotemporal states are effective.
Jieyi Wang, Zeming Liu, Dexuan Xu, Chuan Wang 0002, Ruiyuan Guan, Weihua Yue, Yu Huang 0004
AAAI3
2025 Semi-Supervised Clustering Framework for Fine-grained Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to detect all objects and identify their pairwise relationships existing in the scene. Considering the substantial human labor costs, existing scene graph annotations are often sparse and biased, which result in confusion training with low-frequency predicates. In this work, we design a Semi-Supervised Clustering framework for Scene Graph Generation (SSC-SGG) that uses the sparse labeled data to guide the generation of effective pseudo-labels from unlabeled object pairs, thus enriching the labeled sample space, especially for low-frequency interaction samples. We approach from the perspective of clustering, reducing the problem of confirmation bias in a self-training manner. Specifically, we first enhance the model's robustness to feature extraction via prototype-based clustering, aggregating different relationship augmented features onto the same prototype. Secondly, we design a dynamic pseudo-label assignment algorithm based on a mini-batch, which adjusts the detection sensitivity to different frequency samples from the historical assignment. Finally, we conduct joint training on the pseudo-labels and the labeled data. We conduct experiments on various SGG models and achieve substantial overall performance improvements, demonstrating the effectiveness of SSC-SGG.
Chuan Wang 0002, Shuyi Wu, Jinjing Zhao, Zeming Liu, Liang Yang 0002
AAAI6
2025 ReFF: Reinforcing Format Faithfulness in Language Models Across Varied Tasks
abstract
Following formatting instructions to generate well-structured content is a fundamental yet often unmet capability for large language models (LLMs). To study this capability, which we refer to as format faithfulness, we present FormatBench, a comprehensive format-related benchmark. Compared to previous format-related benchmarks, FormatBench involves a greater variety of tasks in terms of application scenes (traditional NLP tasks, creative works, autonomous agency tasks), human-LLM interaction styles (single-turn instruction, multi-turn chat), and format types (inclusion, wrapping, length, coding). Moreover, each task in FormatBench is attached with a format checker program. Extensive experiments on the benchmark reveal that state-of-the-art open- and closed-source LLMs still suffer from severe deficiency in format faithfulness. By virtue of the decidable nature of formats, we propose to Reinforce Format Faithfulness (ReFF) to help LLMs generate formatted output as instructed without compromising general quality. Without any annotated data, ReFF can substantially improve the format faithfulness rate (e.g., from 21.6% in original LLaMA3 to 95.0% on caption segmentation task), while keep the general quality comparable (e.g., from 47.3 to 46.4 in F1 scores). Combined with labeled training data, ReFF can simultaneously improve both format faithfulness (e.g., from 21.6% in original LLaMA3 to 75.5%) and general quality (e.g., from 47.3 to 61.6 in F1 scores). We further offer an interpretability analysis to explain how ReFF improves both format faithfulness and general quality.
Jiashu Yao, Heyan Huang, Zeming Liu, Haoyu Wen, Boao Qian, Yuhang Guo 0001
AAAI3
2025 GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
abstract
Video Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural and contextual subtleties. Although Multimodal Large Language Models (MLLMs) and Chain-of-Thought (CoT) have demonstrated strong reasoning abilities in STEM tasks (e.g. mathematics and coding), they still struggle to generate creative expressions such as resonant jokes and insightful satire. Moreover, existing benchmarks are constrained by their limited modalities and insufficient categories, hindering the exploration of comprehensive creativity in video-based Comment Art creation. To address these limitations, we introduce GODBench, a novel benchmark that integrates video and text modalities to systematically evaluate MLLMs’ abilities to compose Comment Art. Furthermore, inspired by the propagation patterns of waves in physics, we propose Ripple of Thought (RoT), a multi-step reasoning framework designed to enhance the creativity of MLLMs. Extensive experiments on GODBench reveal that existing MLLMs and CoT methods still face significant challenges in understanding and generating creative video comments. In contrast, RoT provides an effective approach to improving creative composing, highlighting its potential to drive meaningful advancements in MLLM-based creativity.
Yiming Lei 0001, Zeming Liu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu 0001, Yunhong Wang 0001
ACL (1)3
2025 HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
abstract
Large language models (LLMs) have the potential to revolutionize smart home assistants by enhancing their ability to accurately understand user needs and respond appropriately, which is extremely beneficial for building a smarter home environment.While recent studies have explored integrating LLMs into smart home systems, they primarily focus on handling straightforward, valid single-device operation instructions.However, real-world scenarios are far more complex and often involve users issuing invalid instructions or controlling multiple devices simultaneously.These have two main challenges: LLMs must accurately identify and rectify errors in user instructions and execute multiple user instructions perfectly.To address these challenges and advance the development of LLM-based smart home assistants, we introduce HomeBench, the first smart home dataset with valid and invalid instructions across single and multiple devices in this paper.We have experimental results on 13 distinct LLMs; e.g., GPT-4o achieves only a 0.0% success rate in the scenario of invalid multi-device instructions, revealing that the existing stateof-the-art LLMs still cannot perform well in this situation even with the help of in-context learning, retrieval-augmented generation, and fine-tuning.Our code and dataset are publicly available at the link 1 .
Silin Li, Yuhang Guo 0001, Jiashu Yao, Zeming Liu, Haifeng Wang 0001
ACL (1)4
2025 Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
abstract
Jiayi Zeng, Yizhe Feng, Mengliang He, Wenhui Lei, Wei Zhang, Zeming Liu, Xiaoming Shi, Aimin Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yizhe Feng, Mengliang He, Wenhui Lei, Zeming Liu, Aimin Zhou
ACL (1)6
2025 SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
abstract
With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities of these models. However, these benchmarks focus on standalone videos and only assess "visual elements" like human actions and object states. In reality, contemporary videos often encompass complex and continuous narratives, typically presented as a series. To address this challenge, we propose SeriesBench, a benchmark consisting of 105 carefully curated narrative-driven series, covering 28 specialized tasks that require deep narrative understanding to solve. Specifically, we first select a diverse set of drama series spanning various genres. Then, we introduce a novel long-span narrative annotation method, combined with a full-information transformation approach to convert manual annotations into diverse task formats. To further enhance the model’s capacity for detailed analysis of plot structures and character relationships within series, we propose a novel narrative reasoning framework, PC-DCoT. Extensive results on SeriesBench indicate that existing MLLMs still face significant challenges in understanding narrative-driven series, while PC-DCoT enables these MLLMs to achieve performance improvements. Overall, our SeriesBench and PC-DCoT highlight the critical necessity of advancing model capabilities for understanding narrative-driven series, guiding future MLLMs development. SeriesBench is publicly available at https://github.com/zackhxn/SeriesBench-CVPR2025.
Yiming Lei 0001, Zeming Liu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu 0001, Yunhong Wang 0001
CVPR3
2025 RETAIL: Towards Real-world Travel Planning for Large Language Models
abstract
Although large language models have enhanced automated travel planning abilities, current systems remain misaligned with real-world scenarios.First, they assume users provide explicit queries, while in reality requirements are often implicit.Second, existing solutions ignore diverse environmental factors and user preferences, limiting the feasibility of plans.Third, systems can only generate plans with basic POI arrangements, failing to provide all-in-one plans with rich details.To mitigate these challenges, we construct a novel dataset RETAIL, which supports decision-making for implicit queries while covering explicit queries, both with and without revision needs.It also enables environmental awareness to ensure plan feasibility under real-world scenarios, while incorporating detailed POI information for allin-one travel plans.Furthermore, we propose a topic-guided multi-agent framework, termed TGMA.Our experiments reveal that even the strongest existing model achieves merely a 1.0% pass rate, indicating real-world travel planning remains extremely challenging.In contrast, TGMA demonstrates substantially improved performance 2.72%, offering promising directions for real-world travel planning.1 sistant for personalized travel planning.
Yizhe Feng, Zeming Liu, Xiangrong Zhu 0002, Yuanfang Guo, Yunhong Wang 0001
EMNLP3
2025 PRIM: Towards Practical In-Image Multilingual Machine Translation
abstract
In-Image Machine Translation (IIMT) aims to translate images containing texts from one language to another. Current research of end-to-end IIMT mainly conducts on synthetic data, with simple background, single font, fixed text position, and bilingual translation, which can not fully reflect real world, causing a significant gap between the research and practical conditions. To facilitate research of IIMT in real-world scenarios, we explore Practical In-Image Multilingual Machine Translation (IIMMT). In order to convince the lack of publicly available data, we annotate the PRIM dataset, which contains real-world captured one-line text images with complex background, various fonts, diverse text positions, and supports multilingual translation directions. We propose an end-to-end model VisTrans to handle the challenge of practical conditions in PRIM, which processes visual text and background information in the image separately, ensuring the capability of multilingual translation while improving the visual quality. Experimental results indicate the VisTrans achieves a better translation quality and visual effect compared to other models. The code and dataset are available at: https://github.com/BITHLP/PRIM.
Yanzhi Tian, Zeming Liu, Chong Feng 0001, Heyan Huang, Yuhang Guo 0001
EMNLP2
2025 Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring
abstract
Honglin Mu, Han He, Yuxin Zhou, Yunlong Feng, Yang Xu, Libo Qin, Xiaoming Shi, Zeming Liu, Xudong Han, Qi Shi, Qingfu Zhu, Wanxiang Che. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Honglin Mu, Han He, Yunlong Feng, Yang Xu 0049, Libo Qin 0001, Zeming Liu, Qi Shi 0002, Qingfu Zhu, Wanxiang Che
NAACL (Long Papers)8
2025 FINet: Detecting Camouflaged Objects Using Frequency-Aware Information
Chengfeng Zhu, Zeming Liu, Qing Zhang 0004, Xuefeng Wu
PRCV (18)2
2025 Towards few-shot mixed-type dialogue generation
Zeming Liu, Haifeng Wang 0001, Zeyang Lei, Zhengyu Niu, Hua Wu 0003, Wanxiang Che
Sci. China Inf. Sci.1
2025 RaNAS: Resource-Aware Neural Architecture Search for Edge Computing
abstract
Neural architecture search (NAS) for edge devices is often time-consuming because of long-latency deploying and testing on edge devices. The ability to accurately predict the computation cost and memory requirement for convolutional neural networks (CNNs) in advance holds substantial value. Existing work primarily relies on analytical models, which can result in high prediction errors. This article proposes a resource-aware NAS (RaNAS) model based on various features. Additionally, a new graph neural network is introduced to predict inference latency and maximum memory requirements for CNNs on edge devices. Experimental results show that, within the error bound of ±1%, RaNAS achieves an accuracy improvement of approximately 8% for inference latency prediction and about 25% for maximum memory occupancy prediction over the state-of-the-art approaches.
Jianhua Gao 0001, Zeming Liu, Yizhuo Wang 0001, Weixing Ji
ACM Trans. Archit. Code Optim.2
2025 Hard-Label Black-Box Adversarial Attacks for Implicit Scene Interactions
Muxue Liang, Chuan Wang 0002, Siyuan Liang 0004, Aishan Liu, Yanan Cao 0006, Qingyong Li, Zeming Liu, Liang Yang 0002, Xiaochun Cao
IEEE Trans. Inf. Forensics Secur.7
2024 Automatic SQL Query Generation from Code Switched Natural Language Questions on Electronic Medical Records
abstract
Electronic Medical Records (EMRs) meticulously document patient information in relational databases, presenting a challenging task for medical professionals to effectively retrieve this data. Natural Language Question to SQL query (NL2SQL), a critical task in natural language processing (NLP), shows promising performance in addressing this challenge. However, existing medical NL2SQL studies often focus on generating SQL queries from monolingual questions in English. None of them have studied medical NL2SQL for the Chinese domain.In Chinese EMRs, the questions are primarily in Chinese, while many medical terms, such as drugs and diseases, are commonly described in English. This phenomenon is generally known as code-switching (CS) in the field. Given that underlying systems are typically monolingual, CS has proven to pose significant accuracy challenges, as observed in many tasks such as Automatic Speech Recognition (ASR) and Machine Translation (MT). However, the potential effects on EMRs NL2SQL have not been explored. In this study, we investigate the CS-NL2SQL problem, focusing on CS in the context of Chinese questions for the NL2SQL task on EMRs. To assess the model’s performance on this task, we construct the first CS-NL2SQL dataset named CS-MIMICSQL in medical domains. We then explore different CS-NL2SQL architectures along two dimensions: cascaded (translation followed by SQL generation) vs. end-to-end. Our results demonstrate that the proposed end-to-end structure can outperform much better than the cascaded baselines.
Jinyin Nie, Zeming Liu, Dong Lei, Yuanfeng Song
BIBM3
2024 TED-EL: A Corpus for Speech Entity Linking
abstract
Speech entity linking amis to recognize mentions from speech and link them to entities in knowledge bases. Previous work on entity linking mainly focuses on visual context and text context. In contrast, speech entity linking focuses on audio context. In this paper, we first propose the speech entity linking task. To facilitate the study of this task, we propose the first speech entity linking dataset, TED-EL. Our corpus is a high-quality, human-annotated, audio, text, and mention-entity pair parallel dataset derived from Technology, Entertainment, Design (TED) talks and includes a wide range of entity types (24 types). Based on TED-EL, we designed two types of models: ranking-based and generative speech entity linking models. We conducted experiments on the TED-EL dataset for both types of models. The results show that the ranking-based models outperform the generative models, achieving an F1 score of 60.68%.
Silin Li, Ruoyu Song 0002, Tianwei Lan, Zeming Liu, Yuhang Guo 0001
LREC/COLING4
2024 AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction
abstract
Large Language Models (LLMs) can interact with the real world by connecting with versatile external APIs, resulting in better problemsolving and task automation capabilities.Previous research primarily focuses on APIs with limited arguments from a single source or overlooks the complex dependency relationship between different APIs.However, it is essential to utilize multiple APIs collaboratively from various sources (e.g., different Apps in the iPhone), especially for complex user instructions.In this paper, we introduce AppBench, the first benchmark to evaluate LLMs' ability to plan and execute multiple APIs from various sources in order to complete the user's task.Specifically, we consider two significant challenges in multiple APIs: 1) graph structures: some APIs can be executed independently while others need to be executed one by one, resulting in graph-like execution order; and 2) permission constraints: which source is authorized to execute the API call.We have experimental results on 9 distinct LLMs; e.g., GPT-4o achieves only a 2.0% success rate at the most complex instruction, revealing that the existing state-of-the-art LLMs still cannot perform well in this situation even with the help of in-context learning and finetuning.
Hongru Wang 0003, Rui Wang 0092, Boyang Xue, Heming Xia, Jingtao Cao, Zeming Liu, Jeff Z. Pan, Kam-Fai Wong
EMNLP6
2024 FAME: Towards Factual Multi-Task Model Editing
abstract
Large language models (LLMs) embed extensive knowledge and utilize it to perform exceptionally well across various tasks.Nevertheless, outdated knowledge or factual errors within LLMs can lead to misleading or incorrect responses, causing significant issues in practical applications.To rectify the fatal flaw without the necessity for costly model retraining, various model editing approaches have been proposed to correct inaccurate knowledge within LLMs in a cost-efficient way.To evaluate these model editing methods, previous work introduced a series of datasets.However, most of the previous datasets only contain fabricated data in a single format, which diverges from realworld model editing scenarios, raising doubts about their usability in practice.To facilitate the application of model editing in real-world scenarios, we propose the challenge of practicality.To resolve such challenges and effectively enhance the capabilities of LLMs, we present FAME, an factual, comprehensive, and multi-task dataset, which is designed to enhance the practicality of model editing.We then propose SKEME, a model editing method that uses a novel caching mechanism to ensure synchronization with the real world.The experiments demonstrate that SKEME performs excellently across various tasks and scenarios, confirming its practicality.
Zeng Li 0002, Yingyu Shan, Zeming Liu, Jiashu Yao, Yuhang Guo 0001
EMNLP3
2024 Bi-directional Boundary-object interaction and refinement network for Camouflaged Object Detection
abstract
Due to the high intrinsic similarity between the camouflaged objects and the background, the predicted edge cue might be inaccurate or even erroneous. This will degrade the detection performance when such edge cues are directly integrated with the camouflaged object features. To solve this issue, we propose a bi-directional camouflaged object detection network to progressive complement and correct the boundary feature and the object feature in an interactive learning manner, thereby achieving prediction with fine structures. Specifically, we design a bi-directional transmission and correction (BTC) module, which contains the boundary and object interaction (BOI) module and the Feature Calibration Fusion (FCF) module, to explore the correlation between the boundary feature and the object feature and then provide the important complementary cues to refine each other, thereby progressively compensating the deficiencies and correcting the mistakes. Experimental results on four benchmark datasets demonstrate the effectiveness and superiority of our network. Our code is publicly available at: https://github.com/Jcogito/BIRNet.
Jicheng Yang, Qing Zhang 0004, Yuetong Li, Zeming Liu
ICME5
2024 2M-NER: contrastive learning for multilingual and multimodal NER with language and modal fusion
Dongsheng Wang 0007, Xiaoqin Feng, Zeming Liu, Chuan Wang 0002
Appl. Intell.3
2024 Dual-space Hierarchical Learning for Goal-guided Conversational Recommendation
Can Chen 0005, Hao Liu 0026, Zeming Liu, Xue (Steve) Liu, Dejing Dou
Neurocomputing3
2023 XDailyDialog: A Multilingual Parallel Dialogue Corpus
abstract
Zeming Liu, Ping Nie, Jie Cai, Haifeng Wang, Zheng-Yu Niu, Peng Zhang, Mrinmaya Sachan, Kaiping Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zeming Liu, Ping Nie, Haifeng Wang 0001, Zhengyu Niu, Mrinmaya Sachan, Kaiping Peng
ACL (1)1
2023 MidMed: Towards Mixed-Type Dialogues for Medical Consultation
abstract
Xiaoming Shi, Zeming Liu, Chuan Wang, Haitao Leng, Kui Xue, Xiaofan Zhang, Shaoting Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zeming Liu, Chuan Wang 0002, Haitao Leng, Kui Xue, Xiaofan Zhang 0002, Shaoting Zhang 0001
ACL (1)2
2023 Focusing on Flexible Masks: A Novel Framework for Panoptic Scene Graph Generation with Relation Constraints
abstract
Panoptic Scene Graph Generation (PSG) presents pixel-wise instance detection and localization, leading to comprehensive and precise scene graphs. Current methods employ conventional Scene Graph Generation (SGG) frameworks to solve the PSG problem, neglecting the fundamental differences between bounding boxes and masks, i.e., bounding boxes are allowed overlap but masks are not. Since segmentation from the panoptic head has deviations, non-overlapping masks may not afford complete instance information. Subsequently, in the training phase, incomplete segmented instances may not be well-aligned to annotated ones, causing mismatched relations and insufficient training. During the inference phase, incomplete segmentation leads to incomplete scene graph prediction. To alleviate these problems, we construct a novel two-stage framework for the PSG problem. In the training phase, we design a proposal matching strategy, which replaces deterministic segmentation results with proposals extracted from the off-the-shelf panoptic head for label alignment, thereby ensuring the all-matching of training samples. In the inference phase, we present an innovative concept of employing relation predictions to constrain segmentation and design a relation-constrained segmentation algorithm. By reconstructing the process of generating segmentation results from proposals using predicted relation results, the algorithm recovers more valid instances and predicts more complete scene graphs. The experimental results show overall superiority, effectiveness, and robustness against adversarial attacks.
Chuan Wang 0002, Zeming Liu, Jiahong Wu 0002, Dongsheng Wang 0007, Liang Yang 0002, Xiaochun Cao
ACM Multimedia3
2023 Graph-Grounded Goal Planning for Conversational Recommendation
abstract
Conversational recommendation casts the recommendation problem as a dialog-based interactive task, which could acquire user interest more efficiently and effectively by allowing users to express what they like. In this work, we move a step towards a new conversational recommendation task that is more suitable for real-world applications. In this task, the recommender proactively and naturally lead a dialog from non-recommendation content to approach an item being of interest to users, and allow users to ask questions for better support of user decisions. The challenge of this task lies in how to effectively control the dialog flow to complete the recommendation while appropriately responding to user utterances. To address this challenge, we first construct a Chinese recommendation dialog dataset DuRecDial. We then propose a two-stage Multi-Goal driven Conversation Generation framework, MGCG. In particular, the goal planning module leverages the global graph structure information and local goal-sequence information to effectively control the dialog flow step by step. The goal-guided responding module can produce an in-depth dialog about each goal by fully exploiting hierarchical goal information for response retrieval or generation. Results on DuRecDial demonstrate that MGCG can lead the dialog more proactively and naturally, and complete the recommendation task more effectively.
Zeming Liu, Hao Liu 0026, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001, Hui Xiong 0001
IEEE Trans. Knowl. Data Eng.1
2022 Where to Go for the Holidays: Towards Mixed-Type Dialogs for Clarification of User Goals
abstract
Most dialog systems posit that users have figured out clear and specific goals before starting an interaction.For example, users have determined the departure, the destination, and the travel time for booking a flight.However, in many scenarios, limited by experience and knowledge, users may know what they need, but still struggle to figure out clear and specific goals by determining all the necessary slots.
Zeming Liu, Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003
ACL (1)1
2021 DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational Recommendation
abstract
In this paper, we provide a bilingual parallel human-to-human recommendation dialog dataset (DuRecDial 2.0) to enable researchers to explore a challenging task of multilingual and cross-lingual conversational recommendation.The difference between DuRecDial 2.0 and existing conversational recommendation datasets is that the data item (Profile, Goal, Knowledge, Context, Response) in DuRecDial 2.0 is annotated in two languages, both English and Chinese, while other datasets are built with the setting of a single language.We collect 8.2k dialogs aligned across English and Chinese languages (16.5k dialogs and 255k utterances in total) that are annotated by crowdsourced workers with strict quality control procedure.We then build monolingual, multilingual, and cross-lingual conversational recommendation baselines on DuRecDial 2.0.Experiment results show that the use of additional English data can bring performance improvement for Chinese conversational recommendation, indicating the benefits of DuRecDial 2.0.Finally, this dataset provides a challenging testbed for future studies of monolingual, multilingual, and cross-lingual conversational recommendation. 1
Zeming Liu, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che
EMNLP (1)1
2020 Towards Conversational Recommendation over Multi-Type Dialogs
abstract
We focus on the study of conversational recommendation in the context of multi-type dialogs, where the bots can proactively and naturally lead a conversation from a nonrecommendation dialog (e.g., QA) to a recommendation dialog, taking into account user's interests and feedback.To facilitate the study of this task, we create a human-to-human Chinese dialog dataset DuRecDial (about 10k dialogs, 156k utterances), which contains multiple sequential dialogs for every pair of a recommendation seeker (user) and a recommender (bot).In each dialog, the recommender proactively leads a multi-type dialog to approach recommendation targets and then makes multiple recommendations with rich interaction behavior.This dataset allows us to systematically investigate different parts of the overall problem, e.g., how to naturally lead a dialog, how to interact with users for recommendation.Finally we establish baseline results on DuRecDial for future studies. 1
Zeming Liu, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001
ACL1
2019 Detection of Cell Types from Single-cell RNA-seq Data using Similarity via Kernel Preserving Learning Embedding
abstract
The recent advances in single-cell sequencing techniques allow us to study biological issues on cell levels. Detecting cell types from scRNA-seq data analysis is important and meaningful. However, high-level noise and the nonlinearity and sparsity of scRNA-seq data are great challenges. In this paper, we propose a cell-type detection algorithm preserving the overall cell relations named POCR to analyze scRNA-seq data. POCR utilizes a kernel embedding similarity measure to calculate cell-to-cell similarity, by minimizing the reconstruction error of a kernel matrix, rather than the reconstruction error of the original data adopted by other similarity metrics. According to the scale of scRNA-seq datasets, we select Gaussian kernel or linear kernel to calculate the embedding. We then adopt spectral clustering to detect the cell types based on the learned cell-to-cell similarity. The results are further visualized to demonstrate the effectiveness of the cell-type detection algorithm POCR. Further analysis shows that the learned similarity could improve the clustering and visualization of cell types in scRNA-seq data. Our proposed algorithm is compared with five other state-of-the-art cell subtype detection methods. The effectiveness of the algorithms is evaluated by two criteria: ARI and NMI. The experiments show that POCR achieves accurate and robust performance across different scRNA-seq data. Our python implementation of POCR is available at https://github.com/ZeMing-Liu/POCR.
Zeming Liu, Chengzhi Hong, Yi-Ping Phoebe Chen, Shichao Liu 0002, Wen Zhang 0008
BIBM1