Biyang Ma

dblp:163/4075 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-1515-6449ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Contrastive generative learning for enhanced multiagent decision-making under uncertainty
abstract
This paper investigates the complexity of intelligent decision-making within multiagent systems operating under uncertainty, particularly emphasizing the challenges associated with modeling behaviors of other agents and optimizing decision-making for a subject agent in a common environment characterized by incomplete historical data. To address these challenges, we propose a generative learning method based on a general multiagent decision making framework, namely interactive dynamic influence diagrams, and apply contrastive learning to diversify the generation of potential behaviors, thereby enhancing the subject agent’s modeling and prediction capabilities. We conduct experiments on multiple classic domains to demonstrate the efficacy of the new learning method in improving decision-making quality. The empirical results highlight its substantial improvement in enhancing overall performance in multiagent decision-making. Our work contributes to multiagent decision making particularly when a subject agent interacts with other unknown agents, including humans, in many practical applications.
Yinghui Pan, Xinyi Xiang, Yifeng Zeng, Biyang Ma, Guoquan Liu, Yew-Soon Ong
Eng. Appl. Artif. Intell.4
2025 Active legibility in multiagent reinforcement learning
abstract
A multiagent sequential decision problem has been seen in many critical applications including urban transportation, autonomous driving cars, military operations, etc. Its widely known solution, namely multiagent reinforcement learning, has evolved tremendously in recent years. Among them, the solution paradigm of modeling other agents attracts our interest, which is different from traditional value decomposition or communication mechanisms. It enables agents to understand and anticipate others' behaviors and facilitates their collaboration. Inspired by recent research on the legibility that allows agents to reveal their intentions through their behavior, we propose a multiagent active legibility framework to improve their performance. The legibility-oriented framework drives agents to conduct legible actions so as to help others optimise their behaviors. In addition, we design a series of problem domains that emulate a common legibility-needed scenario and effectively characterize the legibility in multiagent reinforcement learning. The experimental results demonstrate that the new framework is more efficient and requires less training time compared to several multiagent reinforcement learning algorithms.
Yanyu Liu, Yinghui Pan, Yifeng Zeng, Biyang Ma, Prashant Doshi
Artif. Intell.4
2025 Multi-population differential evolution approach for feature selection with mutual information ranking
Fei Yu 0008, Hongrun Wu, Hui Wang 0026, Biyang Ma
Expert Syst. Appl.5
2023 Symmetric Bayesian Personalized Ranking With Softmax Weight
abstract
Preference learning, especially pairwise preference learning, is an efficient method for modeling implicit feedback in item recommendation. However, it is insufficient and not always valid to work for a basic pairwise preference model, which assumes that users prefer interacted (i.e., bought or viewed) items to un-interacted (i.e., not bought or not viewed) items. Recently, the state-of-the-art approaches have emerged as two separate but powerful methods, namely, pairwise preferences over item-sets and asymmetric pairwise preference models, respectively, to address limitations of the basic models. In spite of the success achieved by these methods, the assumption that the horizontal pairwise preference is with respect to two items does not always hold in asymmetric pairwise preference. Hence, it is appealing to integrate them into a uniform approach. In this article, we propose a novel symmetric pairwise preference assumption. We use a weighted average through a softmax function and define the overall preferences that can better discover users’ preference patterns. With the new assumption and the weighted average method, we propose a novel recommendation algorithm to improve the recommendation quality. Extensive empirical studies show that our new algorithms can significantly outperform several state-of-the-art and baseline methods over a number of public datasets.
Yinghui Pan, Qiang Ran, Yifeng Zeng, Biyang Ma, Jing Tang 0001, Langcai Cao
IEEE Trans. Syst. Man Cybern. Syst.4
2022 Tensor decomposition for multi-agent predictive state representation
Biyang Ma, Bilian Chen, Yifeng Zeng, Jing Tang 0001, Langcai Cao
Expert Syst. Appl.1
2022 Diversifying agent's behaviors in interactive decision models
abstract
Modeling other agents' behaviors plays an important role in decision models for interactions among multiple agents. To optimize its own decisions, a subject agent needs to model what other agents act simultaneously in an uncertain environment. However, modeling insufficiency occurs when the agents are competitive and the subject agent cannot get full knowledge about other agents. Even when the agents are collaborative, they may not share their true behaviors due to their privacy concerns. Most of the recent research still assumes that the agents have common knowledge about their environments and a subject agent has the true behavior of other agents in its mind. Consequently, the resulting techniques are not applicable in many practical problem domains. In this article, we investigate into diversifying behaviors of other agents in the subject agent's decision model before their interactions. The challenges lie in generating and measuring new behaviors of other agents. Starting with prior knowledge about other agents' behaviors, we use a linear reduction technique to extract representative behavioral features from the known behaviors. We subsequently generate their new behaviors by expanding the features and propose two diversity measurements to select top- K $K$ behaviors. We demonstrate the performance of the new techniques in two well-studied problem domains. The top- K $K$ behavior selection embarks the study of unknown behaviors in multiagent decision making and inspires investigation of diversifying agents' behaviors in competitive agent interactions. This study will contribute to intelligent systems dealing with unknown unknowns in an open artificial intelligence world.
Yinghui Pan, Hanyi Zhang, Yifeng Zeng, Biyang Ma, Jing Tang 0001, Zhong Ming 0001
Int. J. Intell. Syst.4
2022 Behavioral model summarisation for other agents under uncertainty
Yinghui Pan, Biyang Ma, Jing Tang 0001, Yifeng Zeng
Inf. Sci.2
2021 Toward data-driven solutions to interactive dynamic influence diagrams
abstract
Abstract With the availability of significant amount of data, data-driven decision making becomes an alternative way for solving complex multiagent decision problems. Instead of using domain knowledge to explicitly build decision models, the data-driven approach learns decisions (probably optimal ones) from available data. This removes the knowledge bottleneck in the traditional knowledge-driven decision making, which requires a strong support from domain experts. In this paper, we study data-driven decision making in the context of interactive dynamic influence diagrams (I-DIDs)—a general framework for multiagent sequential decision making under uncertainty. We propose a data-driven framework to solve the I-DIDs model and focus on learning the behavior of other agents in problem domains. The challenge is on learning a complete policy tree that will be embedded in the I-DIDs models due to limited data. We propose two new methods to develop complete policy trees for the other agents in the I-DIDs. The first method uses a simple clustering process, while the second one employs sophisticated statistical checks. We analyze the proposed algorithms in a theoretical way and experiment them over two problem domains.
Yinghui Pan, Jing Tang 0001, Biyang Ma, Yifeng Zeng, Zhong Ming 0001
Knowl. Inf. Syst.3
2021 Tensor optimization with group lasso for multi-agent predictive state representation
Biyang Ma, Jing Tang 0001, Bilian Chen, Yinghui Pan, Yifeng Zeng
Knowl. Based Syst.1
2017 Group sparse optimization for learning predictive state representations
Yifeng Zeng, Biyang Ma, Bilian Chen, Jing Tang 0001, Mengda He
Inf. Sci.2