Mingfei Sun 0001

dblp:195/7934 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Accelerating LLM-Based Algorithm Evolution for the 3D Container Loading Problem
abstract
Designing effective heuristics remains a labor-intensive task traditionally reserved for domain experts. While Large Language Model (LLM)-driven evolutionary search offers a path toward automated discovery, existing methods often suffer from slow convergence, primarily due to inefficient hyperparameter tuning. Delegating tuning to specialized optimizers improves performance, but it comes at the cost of code bloat and overfitting. To address these issues, we propose a pipeline that introduces a novel regularization architecture balancing performance and complexity. Specifically, we mitigate the side effects of automated tuning through two novel components: (i) symbolic pruning mutator, which combines LLM semantic guidance with Abstract Syntax Tree analysis to eliminate algorithmic redundancy; and (ii) a complexity-aware mutation gate that explicitly filters out mutations leading to excessive code growth. Our framework substantially accelerates convergence and improves generalization on standard benchmarks of the 3D Single Container Loading Problem. The discovered heuristics match state-of-the-art human designed algorithms and rediscover similar geometric principles used by experts, highlighting the framework's ability to autonomously extract meaningful domain knowledge.
Guorui Quan, Mingfei Sun 0001, Manuel López-Ibáñez 0001, Nicolás Álvarez-Gil, Silvino Fernandez
GECCO2
2025 Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman Filtering
abstract
Vision-language pretrained models, such as CLIP, have established new benchmarks in multimodal data mining. In such models, few-shot fine-tuning is a major challenge to achieve optimal performance on both in-distribution (ID) and out-of-distribution (OOD) datasets, especially when labeled data is scarce. Most existing fine-tuning approaches rely on first-order gradient-based optimizers, which typically suffer from slow convergence, sensitivity to step-size hyperparameters, and poor generalization in OOD settings. In contrast, second-order methods utilize local curvature information of the loss landscape to adjust the update step size. This is particularly beneficial for CLIP models, whose non-convex loss functions often contain sharp critical points. In such cases, natural gradient direction can offer more substantial and efficient per-iteration updates when fine-tuning with limited data. Natural Gradient Descent (NGD) is obtained by preconditioning the standard gradient with the inverse Fisher Information Matrix (FIM), which is computationally expensive for large models. To address this, we propose a Bayesian approximation of NGD using a Kalman filter for CLIP models. Our method combines the benefits of second-order optimization with Bayesian inference, which enhances generalization while providing uncertainty quantification. Extensive experiments conducted on diverse image classification datasets demonstrate that our algorithm consistently achieves superior-or comparable-ID performance and improved OOD robustness compared to state-of-the-art baselines. To the best of our knowledge, this work represents the first successful application of Kalman filtering to fine-tuning CLIP-based models, which enables more robust and efficient learning in vision-language tasks.
Hossein Abdi, Mingfei Sun 0001, Wei Pan 0004
ICDM2
2025 DroneDiffusion: Robust Quadrotor Dynamics Learning with Diffusion Models
abstract
An inherent fragility of quadrotor systems stems from model inaccuracies and external disturbances. These factors hinder performance and compromise the stability of the system, making precise control challenging. Existing model-based approaches either make deterministic assumptions, utilize Gaussian-based representations of uncertainty, or rely on nominal models, all of which often fall short in capturing the complex, multimodal nature of real-world dynamics. This work introduces DroneDiffusion, a novel framework that leverages conditional diffusion models to learn quadrotor dynamics, formulated as a sequence generation task. DroneDiffusion achieves superior generalization to unseen, complex scenarios by capturing the temporal nature of uncertainties and mitigating error propagation. We integrate the learned dynamics with an adaptive controller for trajectory tracking with stability guarantees. Extensive experiments in both simulation and real-world flights demonstrate the robustness of the framework across a range of scenarios, including unfamiliar flight paths and varying payloads, velocities, and wind disturbances. Project page: https://sites.google.com/view/dronediffusion.
Avirup Das, Rishabh Dev Yadav, Sihao Sun, Mingfei Sun 0001, Samuel Kaski, Wei Pan 0004
ICRA4
2024 FARPLS: A Feature-Augmented Robot Trajectory Preference Labeling System to Assist Human Labelers' Preference Elicitation
abstract
Preference-based learning aims to align robot task objectives with human values. One of the most common methods to infer human preferences is by pairwise comparisons of robot task trajectories. Traditional comparison-based preference labeling systems seldom support labelers to digest and identify critical differences between complex trajectories recorded in videos. Our formative study (N = 12) suggests that individuals may overlook non-salient task features and establish biased preference criteria during their preference elicitation process because of partial observations. In addition, they may experience mental fatigue when given many pairs to compare, causing their label quality to deteriorate. To mitigate these issues, we propose FARPLS, a Feature-Augmented Robot trajectory Preference Labeling System. FARPLS highlights potential outliers in a wide variety of task features that matter to humans and extracts the corresponding video keyframes for easy review and comparison. It also dynamically adjusts the labeling order according to users’ familiarities, difficulties of the trajectory pair, and level of disagreements. At the same time, the system monitors labelers’ consistency and provides feedback on labeling progress to keep labelers engaged. A between-subjects study (N = 42, 105 pairs of robot pick-and-place trajectories per person) shows that FARPLS can help users establish preference criteria more easily and notice more relevant details in the presented trajectories than the conventional interface. FARPLS also improves labeling consistency and engagement, mitigating challenges in preference elicitation without raising cognitive loads significantly.
Hanfang Lyu, Yuanchen Bai, Ujaan Das, Chuhan Shi, Leiliang Gong, Yingchi Li, Mingfei Sun 0001, Ming Ge, Xiaojuan Ma
IUI8
2024 Effective Generation of Feasible Solutions for Integer Programming via Guided Diffusion
abstract
Feasible solutions are crucial for Integer Programming (IP) since they can substantially speed up the solving process. In many applications, similar IP instances often exhibit similar structures and shared solution distributions, which can be potentially modeled by deep learning methods. Unfortunately, existing deep-learning-based algorithms, such as Neural Diving [21] and Predict-and-search framework [8], are limited to generating only partial feasible solutions, and they must rely on solvers like SCIP and Gurobi to complete the solutions for a given IP problem. In this paper, we propose a novel framework that generates complete feasible solutions end-to-end. Our framework leverages contrastive learning to characterize the relationship between IP instances and solutions, and learns latent embeddings for both IP instances and their solutions. Further, the framework employs diffusion models to learn the distribution of solution embeddings conditioned on IP representations, with a dedicated guided sampling strategy that accounts for both constraints and objectives. We empirically evaluate our framework on four typical datasets of IP problems, and show that it effectively generates complete feasible solutions with a high probability (> 89.7 %) without the reliance of Solvers and the quality of solutions is comparable to the best heuristic solutions from Gurobi. Furthermore, by integrating our method's sampled partial solutions with the CompleteSol heuristic from SCIP [19], the resulting feasible solutions outperform those from state-of-the-art methods across all datasets, exhibiting a 3.7 to 33.7% improvement in the gap to optimal values, and maintaining a feasible ratio of over 99.7% for all datasets.
Avirup Das, Junying He, Kunpeng Han, Haoyuan Hu, Mingfei Sun 0001
KDD7
2023 Imitating Human Behaviour with Diffusion Models
Tim Pearce, Tabish Rashid, Anssi Kanervisto, David Bignell, Mingfei Sun 0001, Raluca Georgescu, Sergio Valcarcel Macua, Shan Zheng Tan, Ida Momennejad, Katja Hofmann, Sam Devlin
ICLR5
2023 SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning
abstract
The availability of challenging benchmarks has played a key role in the recent progress of machine learning. In cooperative multi-agent reinforcement learning, the StarCraft Multi-Agent Challenge (SMAC) has become a popular testbed for centralised training with decentralised execution. However, after years of sustained improvement on SMAC, algorithms now achieve near-perfect performance. In this work, we conduct new analysis demonstrating that SMAC lacks the stochasticity and partial observability to require complex closed-loop policies. In particular, we show that an open-loop policy conditioned only on the timestep can achieve non-trivial win rates for many SMAC scenarios. To address this limitation, we introduce SMACv2, a new version of the benchmark where scenarios are procedurally generated and require agents to generalise to previously unseen settings (from the same distribution) during evaluation. We also introduce the extended partial observability challenge (EPO), which augments SMACv2 to ensure meaningful partial observability. We show that these changes ensure the benchmarkrequires the use of closed-loop policies. We evaluate state-of-the-art algorithms on SMACv2 and show that it presents significant challenges not present in the original benchmark. Our analysis illustrates that SMACv2 addresses the discovered deficiencies of SMAC and can help benchmark the next generation of MARL methods. Videos of training are available on our website.
Benjamin Ellis, Jonathan Cook 0004, Skander Moalla, Mikayel Samvelyan, Mingfei Sun 0001, Anuj Mahajan, Jakob N. Foerster, Shimon Whiteson
NeurIPS5
2023 CrowdPatrol: A Mobile Crowdsensing Framework for Traffic Violation Hotspot Patrolling
abstract
Traffic violations have become one of the major threats to urban transportation systems, undermining human safety and causing economic losses. To alleviate this problem, crowd-based patrol forces including traffic police and voluntary participants have been employed in many cities. To adaptively optimize patrol routes with limited manpower, it is essential to be aware of traffic violation hotspots. Traditionally, traffic violation hotspots are directly inferred from experiences, and existing patrol routes are usually fixed. In this paper, we propose a mobile crowdsensing-based framework to dynamically infer traffic violation hotspots and adaptively schedule crowd patrol routes. Specifically, we first extract traffic violation-prone locations from heterogeneous crowd-sensed data and propose a spatiotemporal context-aware self-adaptive learning model (CSTA) to infer traffic violation hotspots. Then, we propose a tensor-based integer linear problem modeling method (TILP) to adaptively find optimal patrol routes under human labor constraints. Experiments on real-world data from two Chinese cities (Xiamen and Chengdu) show that our approach accurately infers traffic violation hotspots with F1-scores above 90% in both cities, and generates patrol routes with relative coverage ratios above 85%, significantly outperforming baseline methods.
Zhihan Jiang 0001, Binbin Zhou 0005, Chenhui Lu, Mingfei Sun 0001, Xiaojuan Ma, Xiaoliang Fan, Cheng Wang 0003, Longbiao Chen
IEEE Trans. Mob. Comput.5
2023 Modeling Adaptive Expression of Robot Learning Engagement and Exploring Its Effects on Human Teachers
abstract
Robot Learning from Demonstration (RLfD) allows non-expert users to teach a robot new skills or tasks directly through demonstrations. Although modeled after human–human learning and teaching, existing RLfD methods make robots act as passive observers without the feedback of their learning statuses in the demonstration gathering stage. To facilitate a more transparent teaching process, we propose two mechanisms of Learning Engagement , Z2O-Mode and D2O-Mode, to dynamically adapt robots’ attentional and behavioral engagement expressions to their actual learning status. Through an online user experiment with 48 participants, we find that, compared with two baselines, the two kinds of Learning Engagement can lead to users’ more accurate mental models of the robot’s learning progress, more positive perceptions of the robot, and better teaching experience. Finally, we provide implications for leveraging engagement expression to facilitate transparent human-AI (robot) communication based on our key findings.
Shuai Ma 0005, Mingfei Sun 0001, Xiaojuan Ma
ACM Trans. Comput. Hum. Interact.2
2022 Deterministic and Discriminative Imitation (D2-Imitation): Revisiting Adversarial Imitation for Sample Efficiency
abstract
Sample efficiency is crucial for imitation learning methods to be applicable in real-world applications. Many studies improve sample efficiency by extending adversarial imitation to be off-policy regardless of the fact that these off-policy extensions could either change the original objective or involve complicated optimization. We revisit the foundation of adversarial imitation and propose an off-policy sample efficient approach that requires no adversarial training or min-max optimization. Our formulation capitalizes on two key insights: (1) the similarity between the Bellman equation and the stationary state-action distribution equation allows us to derive a novel temporal difference (TD) learning approach; and (2) the use of a deterministic policy simplifies the TD learning. Combined, these insights yield a practical algorithm, Deterministic and Discriminative Imitation (D2-Imitation), which oper- ates by first partitioning samples into two replay buffers and then learning a deterministic policy via off-policy reinforcement learning. Our empirical results show that D2-Imitation is effective in achieving good sample efficiency, outperforming several off-policy extension approaches of adversarial imitation on many control tasks.
Mingfei Sun 0001, Sam Devlin, Katja Hofmann, Shimon Whiteson
AAAI1
2022 Uni[MASK]: Unified Inference in Sequential Decision Problems
abstract
Randomly masking and predicting word tokens has been a successful approach in pre-training language models for a variety of downstream tasks. In this work, we observe that the same idea also applies naturally to sequential decision making, where many well-studied tasks like behavior cloning, offline RL, inverse dynamics, and waypoint conditioning correspond to different sequence maskings over a sequence of states, actions, and returns. We introduce the UniMASK framework, which provides a unified way to specify models which can be trained on many different sequential decision making tasks. We show that a single UniMASK model is often capable of carrying out many tasks with performance similar to or better than single-task models. Additionally, after fine-tuning, our UniMASK models consistently outperform comparable single-task models.
Micah Carroll, Orr Paradise, Jessy Lin, Raluca Georgescu, Mingfei Sun 0001, David Bignell, Stephanie Milani, Katja Hofmann, Matthew J. Hausknecht, Anca D. Dragan, Sam Devlin
NeurIPS5
2022 Supervised Learning Achieves Human-Level Performance in MOBA Games: A Case Study of Honor of Kings
abstract
We present JueWu-SL, the first supervised-learning-based artificial intelligence (AI) program that achieves human-level performance in playing multiplayer online battle arena (MOBA) games. Unlike prior attempts, we integrate the macro-strategy and the micromanagement of MOBA-game-playing into neural networks in a supervised and end-to-end manner. Tested on Honor of Kings, the most popular MOBA at present, our AI performs competitively at the level of High King players in standard 5v5 games.
Deheng Ye, Peilin Zhao, Fuhao Qiu, Bo Yuan 0008, Mingfei Sun 0001, Siqin Li, Zhenjie Lian, Bei Shi, Liang Wang 0015, Tengfei Shi, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang
IEEE Trans. Neural Networks Learn. Syst.8
2020 Mastering Complex Control in MOBA Games with Deep Reinforcement Learning
abstract
We study the reinforcement learning problem of complex action control in the Multi-player Online Battle Arena (MOBA) 1v1 games. This problem involves far more complicated state and action spaces than those of traditional 1v1 games, such as Go and Atari series, which makes it very difficult to search any policies with human-level performance. In this paper, we present a deep reinforcement learning framework to tackle this problem from the perspectives of both system and algorithm. Our system is of low coupling and high scalability, which enables efficient explorations at large scale. Our algorithm includes several novel strategies, including control dependency decoupling, action mask, target attention, and dual-clip PPO, with which our proposed actor-critic network can be effectively trained in our system. Tested on the MOBA game Honor of Kings, the trained AI agents can defeat top professional human players in full 1v1 games.
Deheng Ye, Mingfei Sun 0001, Bei Shi, Peilin Zhao, Hongsheng Yu, Xipeng Wu, Qingwei Guo, Qiaobo Chen, Yinyuting Yin, Tengfei Shi, Liang Wang 0015, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang
AAAI3
2019 PeerLens: Peer-inspired Interactive Learning Path Planning in Online Question Pool
abstract
Online question pools like LeetCode provide hands-on exercises of skills and knowledge. However, due to the large volume of questions and the intent of hiding the tested knowledge behind them, many users find it hard to decide where to start or how to proceed based on their goals and performance. To overcome these limitations, we present PeerLens, an interactive visual analysis system that enables peer-inspired learning path planning. PeerLens can recommend a customized, adaptable sequence of practice questions to individual learners, based on the exercise history of other users in a similar learning scenario. We propose a new way to model the learning path by submission types and a novel visual design to facilitate the understanding and planning of the learning path. We conducted a within-subject experiment to assess the efficacy and usefulness of PeerLens in comparison with two baseline systems. Experiment results show that users are more confident in arranging their learning path via PeerLens and find it more informative and intuitive.
Meng Xia 0002, Mingfei Sun 0001, Huan Wei, Qing Chen 0001, Yong Wang 0021, Lei Shi 0002, Huamin Qu, Xiaojuan Ma
CHI2
2019 Adversarial Imitation Learning from Incomplete Demonstrations
abstract
Imitation learning targets deriving a mapping from states to actions, a.k.a. policy, from expert demonstrations. Existing methods for imitation learning typically require any actions in the demonstrations to be fully available, which is hard to ensure in real applications. Though algorithms for learning with unobservable actions have been proposed, they focus solely on state information and over- look the fact that the action sequence could still be partially available and provide useful information for policy deriving. In this paper, we propose a novel algorithm called Action-Guided Adversarial Imitation Learning (AGAIL) that learns a pol- icy from demonstrations with incomplete action sequences, i.e., incomplete demonstrations. The core idea of AGAIL is to separate demonstrations into state and action trajectories, and train a policy with state trajectories while using actions as auxiliary information to guide the training whenever applicable. Built upon the Generative Adversarial Imitation Learning, AGAIL has three components: a generator, a discriminator, and a guide. The generator learns a policy with rewards provided by the discriminator, which tries to distinguish state distributions between demonstrations and samples generated by the policy. The guide provides additional rewards to the generator when demonstrated actions for specific states are available. We com- pare AGAIL to other methods on benchmark tasks and show that AGAIL consistently delivers com- parable performance to the state-of-the-art methods even when the action sequence in demonstrations is only partially available.
Mingfei Sun 0001, Xiaojuan Ma
IJCAI1
2017 Sensing and Handling Engagement Dynamics in Human-Robot Interaction Involving Peripheral Computing Devices
abstract
When human partners attend to peripheral computing devices while interacting with conversational robots, the inability of the robots to determine the actual engagement level of the human partners after gaze shift may cause communication breakdown. In this paper, we propose a real-time perception model for robots to estimate human partners' engagement dynamics, and investigate different robot behavior strategies to handle ambiguities in humans' status and ensure the flow of the conversation. In particular, we define four novel types of engagement status and propose a real-time engagement inference model that weighs humans' social signals dynamically according to the involvement of the computing devices. We further design two robot behavior strategies (explicit and implicit) to help resolve uncertainties in engagement inference and mitigate the impact of uncoupling, based on an annotated human-human interaction video corpus. We conducted a within-subject experiment to assess the efficacy and usefulness of the proposed engagement inference model and behavior strategies. Results show that robots with our engagement model can deliver better service and smoother conversations as an assistant, and people find the implicit strategy more polite and appropriate.
Mingfei Sun 0001, Zhenjie Zhao, Xiaojuan Ma
CHI1