Mengdi Li 0006

dblp:183/5712-6 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
7since 2021 · last 2024
0009-0000-2650-2891ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 35% Vision and language · 31% Knowledge representation and reasoning · 21%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
commonsense knowledge acquisition
0.712023
Visually Grounded Commonsense Knowledge Acquisition · AAAI 2023
Machine learning › Reinforcement learning
policy learning
0.712023
Internally Rewarded Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning
reward design
0.712023
Internally Rewarded Reinforcement Learning · ICML 2023
Computer vision › Vision and language
vision-language pretraining
0.712023
Visually Grounded Commonsense Knowledge Acquisition · AAAI 2023
Computer vision › Segmentation and scene understanding
scene graph generation
0.512021
Visual Distant Supervision for Scene Graph Generation · ICCV 2021
Computer vision › Vision and language › visual relationship understanding
visual relation learning
0.512021
Visual Distant Supervision for Scene Graph Generation · ICCV 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
commonsense knowledge base
0.112021
Visual Distant Supervision for Scene Graph Generation · ICCV 2021

Methods — techniques the papers use, named apart from their topics

multi-instance learning · 0.7discriminator-based reward · 0.7contrastive attention · 0.7clipped linear reward function · 0.7semi-supervised learning · 0.5probabilistic label estimation · 0.5distant supervision · 0.5
YearPublicationVenuePosition
2024 Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through Logic
abstract
Recent advancements in large language models have showcased their remarkable generalizability across various domains. However, their reasoning abilities still have significant room for improvement, especially when confronted with scenarios requiring multi-step reasoning. Although large language models possess extensive knowledge, their reasoning often fails to effectively utilize this knowledge to establish a coherent thinking paradigm. These models sometimes show hallucinations as their reasoning procedures are unconstrained by logical principles. Aiming at improving the zero-shot chain-of-thought reasoning ability of large language models, we propose LoT (Logical Thoughts), a self-improvement prompting framework that leverages principles rooted in symbolic logic, particularly Reductio ad Absurdum, to systematically verify and rectify the reasoning processes step by step. Experimental evaluations conducted on language tasks in diverse domains, including arithmetic, commonsense, symbolic, causal inference, and social problems, demonstrate the efficacy of enhanced reasoning by logic. The implementation code for LoT can be accessed at: https://github.com/xf-zhao/LoT.
Xufeng Zhao 0002, Mengdi Li 0006, Wenhao Lu, Cornelius Weber, Jae Hee Lee 0001, Kun Chu, Stefan Wermter
LREC/COLING2
2023 Visually Grounded Commonsense Knowledge Acquisition
abstract
Large-scale commonsense knowledge bases empower a broad range of AI applications, where the automatic extraction of commonsense knowledge (CKE) is a fundamental and challenging problem. CKE from text is known for suffering from the inherent sparsity and reporting bias of commonsense in text. Visual perception, on the other hand, contains rich commonsense knowledge about real-world entities, e.g., (person, can_hold, bottle), which can serve as promising sources for acquiring grounded commonsense knowledge. In this work, we present CLEVER, which formulates CKE as a distantly supervised multi-instance learning problem, where models learn to summarize commonsense relations from a bag of images about an entity pair without any human annotation on image instances. To address the problem, CLEVER leverages vision-language pre-training models for deep understanding of each image in the bag, and selects informative instances from the bag to summarize commonsense entity relations via a novel contrastive attention mechanism. Comprehensive experimental results in held-out and human evaluation show that CLEVER can extract commonsense knowledge in promising quality, outperforming pre-trained language model-based methods by 3.9 AUC and 6.4 mAUC points. The predicted commonsense scores show strong correlation with human judgment with a 0.78 Spearman coefficient. Moreover, the extracted commonsense can also be grounded into images with reasonable interpretability. The data and codes can be obtained at https://github.com/thunlp/CLEVER.
Yuan Yao 0013, Tianyu Yu 0002, Mengdi Li 0006, Ruobing Xie, Cornelius Weber, Zhiyuan Liu 0001, Hai-Tao Zheng 0002, Stefan Wermter, Tat-Seng Chua, Maosong Sun 0001
AAAI4
2023 Internally Rewarded Reinforcement Learning
abstract
We study a class of reinforcement learning problems where the reward signals for policy learning are generated by a discriminator that is dependent on and jointly optimized with the policy. This interdependence between the policy and the discriminator leads to an unstable learning process because reward signals from an immature discriminator are noisy and impede policy learning, and conversely, an under-optimized policy impedes discriminator learning. We call this learning setting $\textit{Internally Rewarded Reinforcement Learning}$ (IRRL) as the reward is not provided directly by the environment but $\textit{internally}$ by the discriminator. In this paper, we formally formulate IRRL and present a class of problems that belong to IRRL. We theoretically derive and empirically analyze the effect of the reward function in IRRL and based on these analyses propose the clipped linear reward function. Experimental results show that the proposed reward function can consistently stabilize the training process by reducing the impact of reward noise, which leads to faster convergence and higher performance compared with baselines in diverse tasks.
Mengdi Li 0006, Xufeng Zhao 0002, Jae Hee Lee 0001, Cornelius Weber, Stefan Wermter
ICML1
2023 Chat with the Environment: Interactive Multimodal Perception Using Large Language Models
abstract
Programming robot behavior in a complex world faces challenges on multiple levels, from dextrous low-level skills to high-level planning and reasoning. Recent pre-trained Large Language Models (LLMs) have shown remarkable reasoning ability in few-shot robotic planning. However, it remains challenging to ground LLMs in multimodal sensory input and continuous action output, while enabling a robot to interact with its environment and acquire novel information as its policies unfold. We develop a robot interaction scenario with a partially observable state, which necessitates a robot to decide on a range of epistemic actions in order to sample sensory information among multiple modalities, before being able to execute the task correctly. An interactive perception framework is therefore proposed with an LLM as its backbone, whose ability is exploited to instruct epistemic actions and to reason over the resulting multimodal sensations (vision, sound, haptics, proprioception), as well as to plan an entire task execution based on the interactively acquired information. Our study demonstrates that LLMs can provide high-level planning and reasoning skills and control interactive robot behavior in a multimodal environment, while multimodal modules with the context of the environmental state help ground the LLMs and extend their processing ability. The project website can be found at https://matcha-model.github.io/.
Xufeng Zhao 0002, Mengdi Li 0006, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter
IROS2
2022 Learning Visually Grounded Human-Robot Dialog in a Hybrid Neural Architecture
Xiaowen Sun, Cornelius Weber, Matthias Kerzel, Tom Weber, Mengdi Li 0006, Stefan Wermter
ICANN (2)5
2021 Visual Distant Supervision for Scene Graph Generation
abstract
Scene graph generation aims to identify objects and their relations in images, providing structured image representations that can facilitate numerous applications in computer vision. However, scene graph models usually require supervised learning on large quantities of labeled data with intensive human annotation. In this work, we propose visual distant supervision, a novel paradigm of visual relation learning, which can train scene graph models without any human-labeled data. The intuition is that by aligning commonsense knowledge bases and images, we can automatically create large-scale labeled data to provide distant supervision for visual relation learning. To alleviate the noise in distantly labeled data, we further propose a framework that iteratively estimates the probabilistic relation labels and eliminates the noisy ones. Comprehensive experimental results show that our distantly supervised model outperforms strong weakly supervised and semi-supervised baselines. By further incorporating human-labeled data in a semi-supervised fashion, our model outperforms state-of-the-art fully supervised models by a large margin (e.g., 8.3 micro- and 7.8 macro-recall@50 improvements for predicate classification in Visual Genome evaluation). We make the data and code for this paper publicly available at https://github.com/thunlp/VisualDS.
Yuan Yao 0011, Xu Han 0007, Mengdi Li 0006, Cornelius Weber, Zhiyuan Liu 0001, Stefan Wermter, Maosong Sun 0001
ICCV4
2021 Robotic Occlusion Reasoning for Efficient Object Existence Prediction
abstract
Reasoning about potential occlusions is essential for robots to efficiently predict whether an object exists in an environment. Though existing work shows that a robot with active perception can achieve various tasks, it is still unclear if occlusion reasoning can be achieved. To answer this question, we introduce the task of robotic object existence prediction: when being asked about an object, a robot needs to move as few steps as possible around a table with randomly placed objects to predict whether the queried object exists. To address this problem, we propose a novel recurrent neural network model that can be jointly trained with supervised and reinforcement learning methods using a curriculum training strategy. Experimental results show that 1) both active perception and occlusion reasoning are necessary to successfully achieve the task; 2) the proposed model demonstrates a good occlusion reasoning ability by achieving a similar prediction accuracy to an exhaustive exploration baseline while requiring only about 10% of the baseline’s number of movement steps on average; and 3) the model generalizes to novel object combinations with a moderate loss of accuracy.
Mengdi Li 0006, Cornelius Weber, Matthias Kerzel, Jae Hee Lee 0001, Zheni Zeng, Zhiyuan Liu 0001, Stefan Wermter
IROS1
2020 Neural Networks for Detecting Irrelevant Questions During Visual Question Answering
Mengdi Li 0006, Cornelius Weber, Stefan Wermter
ICANN (2)1
2019 A novel natural language steganographic framework based on image description neural network
Xuejing Zhou, Mengdi Li 0006, Ping Zhong 0003, Yiming Xue
J. Vis. Commun. Image Represent.3
2019 Generating steganographic image description by dynamic synonym substitution
Mengdi Li 0006, Kai Mu, Ping Zhong 0003, Yiming Xue
Signal Process.1