EDBT 2026 Demo / reviewers in the wild / expert
Minglu Zhao
dblp:284/4795
· DBLP profile ↗
17ranked-venue papers
7as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A minimalistic representation model for head direction system
Minglu Zhao, Dehong Xu, Deqian Kong, Wenhao Zhang 0002, Ying Nian Wu |
CogSci | 1 |
| 2025 | Inverse Attention Agents for Multi-Agent SystemsabstractA major challenge for Multi-Agent Systems (MAS) is enabling agents to adapt dynamically to diverse environments in which opponents and teammates may continually change. Agents trained using conventional methods tend to excel only within the confines of their training cohorts; their performance drops significantly when confronting unfamiliar agents. To address this shortcoming, we introduce Inverse Attention Agents that adopt concepts from the Theory of Mind (ToM) implemented algorithmically using an attention mechanism trained in an end-to-end manner. Crucial to determining the final actions of these agents, the weights in their attention model explicitly represent attention to different goals. We furthermore propose an inverse attention network that deduces the ToM of agents based on observations and prior actions. The network infers the attentional states of other agents, thereby refining the attention weights to adjust the agent's final action. We conduct experiments in a continuous environment, tackling demanding tasks encompassing cooperation, competition, and a blend of both. They demonstrate that the inverse attention network successfully infers the attention of other agents, and that this information improves agent performance. Additional human experiments show that, compared to baseline agent models, our inverse attention agents exhibit superior cooperation with humans and better emulate human behaviors. Qian Long, Ruoyan Li, Minglu Zhao, Demetri Terzopoulos |
ICLR | 3 |
| 2025 | Latent Thought Models with Variational Bayes Inference-Time ComputationabstractWe propose a novel class of language models, Latent Thought Models (LTMs), which incorporate explicit latent thought vectors that follow an explicit prior model in latent space. These latent thought vectors guide the autoregressive generation of ground tokens through a Transformer decoder. Training employs a dual-rate optimization process within the classical variational Bayes framework: fast learning of local variational parameters for the posterior distribution of latent vectors (inference-time computation), and slow learning of global decoder parameters. Empirical studies reveal that LTMs possess additional scaling dimensions beyond traditional Large Language Models (LLMs), such as the number of iterations in inference-time computation and number of latent thought vectors. Higher sample efficiency can be achieved by increasing training compute per token, with further gains possible by trading model size for more inference steps. Designed based on these scaling properties, LTMs demonstrate superior sample and parameter efficiency compared to autoregressive models and discrete diffusion models. They significantly outperform these counterparts in validation perplexity and zero-shot language modeling tasks. Additionally, LTMs exhibit emergent few-shot in-context reasoning capabilities that scale with model size, and achieve competitive performance in conditional and unconditional text generation. The project page is available at https://deqiankong.github.io/blogs/ltm. Deqian Kong, Minglu Zhao, Dehong Xu, Bo Pang 0004, Shu Wang 0002, Edouardo Honig, Zhangzhang Si, Jianwen Xie, Sirui Xie, Ying Nian Wu |
ICML | 2 |
| 2025 | Place Cells as Multi-Scale Position Embeddings: Random Walk Transition Kernels for Path PlanningabstractThe hippocampus supports spatial navigation by encoding cognitive maps through collective place cell activity. We model the place cell population as non-negative spatial embeddings derived from the spectral decomposition of multi-step random walk transition kernels. In this framework, inner product or equivalently Euclidean distance between embeddings encode similarity between locations in terms of their transition probability across multiple scales, forming a cognitive map of adjacency. The combination of non-negativity and inner-product structure naturally induces sparsity, providing a principled explanation for the localized firing fields of place cells without imposing explicit constraints. The temporal parameter that defines the diffusion scale also determines field size, aligning with the hippocampal dorsoventral hierarchy. Our approach constructs global representations efficiently through recursive composition of local transitions, enabling smooth, trap-free navigation and preplay-like trajectory generation. Moreover, theta phase arises intrinsically as the angular relation between embeddings, linking spatial and temporal coding within a single representational geometry. Minglu Zhao, Dehong Xu, Deqian Kong, Wenhao Zhang 0002, Ying Nian Wu |
NeurIPS | 1 |
| 2025 | TIA2V: Video generation conditioned on triple modalities of text-image-audio
Minglu Zhao, Wenmin Wang 0001, Rui Zhang 0108, Haomei Jia |
Expert Syst. Appl. | 1 |
| 2024 | Inverse Reinforcement Learning with Failed Demonstrations towards Stable Driving Behavior ModelingabstractDriving behavior modeling is crucial in autonomous driving systems for preventing traffic accidents. Inverse reinforcement learning (IRL) allows autonomous agents to learn complicated behaviors from expert demonstrations. Similar to how humans learn by trial and error, failed demonstrations can help an agent avoid failures. However, expert and failed demonstrations generally have some common behaviors, which could cause instability in an IRL model. To improve the stability, this work proposes a novel method that introduces time-series labeling for the optimization of IRL to help distinguish the behaviors in demonstrations. Experimental results in a simulated driving environment show that the proposed method converged faster than and outperformed other baseline methods. The results also show consistency for various data balances of the number of expert and failed demonstrations. Minglu Zhao, Masamichi Shimosaka |
IV | 1 |
| 2024 | Latent Plan Transformer for Trajectory Abstraction: Planning as Latent Space InferenceabstractIn tasks aiming for long-term returns, planning becomes essential. We study generative modeling for planning with datasets repurposed from offline reinforcement learning. Specifically, we identify temporal consistency in the absence of step-wise rewards as one key technical challenge. We introduce the Latent Plan Transformer (LPT), a novel model that leverages a latent variable to connect a Transformer- based trajectory generator and the final return. LPT can be learned with maximum likelihood estimation on trajectory-return pairs. In learning, posterior sampling of the latent variable naturally integrates sub-trajectories to form a consistent abstrac- tion despite the finite context. At test time, the latent variable is inferred from an expected return before policy execution, realizing the idea of planning as inference. Our experiments demonstrate that LPT can discover improved decisions from sub- optimal trajectories, achieving competitive performance across several benchmarks, including Gym-Mujoco, Franka Kitchen, Maze2D, and Connect Four. It exhibits capabilities in nuanced credit assignments, trajectory stitching, and adaptation to environmental contingencies. These results validate that latent variable inference can be a strong alternative to step-wise reward prompting. Deqian Kong, Dehong Xu, Minglu Zhao, Bo Pang 0004, Jianwen Xie, Andrew Lizarraga, Sirui Xie, Ying Nian Wu |
NeurIPS | 3 |
| 2024 | SgLFT: Semantic-guided Late Fusion Transformer for video corpus moment retrieval
Tongbao Chen, Wenmin Wang 0001, Minglu Zhao, Ruochen Li 0001 |
Neurocomputing | 3 |
| 2024 | Ultrahigh-definition video quality assessment: A new dataset and benchmark
Ruochen Li 0001, Wenmin Wang 0001, Huanqiang Hu, Tongbao Chen, Minglu Zhao |
Neurocomputing | 5 |
| 2024 | TA2V: Text-Audio Guided Video GenerationabstractRecent conditional and unconditional video generation tasks have been accomplished mainly based on generative adversarial network (GAN), diffusion, and autoregressive models. However, in some circumstances, using only one modality cannot provide enough semantic information. Therefore, in this paper, we propose text-audio to video (TA2V) generation, a new task for generating realistic videos from two different guided modalities, text and audio, which has not been explored much thus far. Compared to image generation, video generation is a harder task because of the complexity of processing higher-dimensional data and scarcer suitable datasets, especially for multimodal video generation. To overcome these limitations, (i) we propose the Text&Audio-guided-Video-Maker (TAgVM) model, which consists of two modules: a text-guided video generator and a text&audio-guided video modifier. (ii) This model uses a 3D VQ-GAN to compress high-dimension video data to a low-dimension discrete sequence, followed by an autoregressive model to guide text-conditional generation in the latent space. Then, we apply a text&audio-guided diffusion model to the generated video scenes, providing additional semantic details corresponding to the audio and text. (iii) We introduce a newly produced music performance video dataset, the University of Rochester Multimodal Music Performance with Video-Audio-Text (URMP-VAT), and a landscape dataset, Landscape with Video-Audio-Text (Landscape-VAT), both of which include three modalities (text, audio, and video) that are aligned with each other. The results demonstrate that our model can create videos with satisfactory quality and semantic information. The source code and datasets are available athttps://github.com/Minglu58/TA2V. Minglu Zhao, Wenmin Wang 0001, Tongbao Chen, Rui Zhang 0108, Ruochen Li 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Intentional commitment through an internalized theory of mind: Acting in the eyes of an imagined observer
Shaozhe Cheng, Minglu Zhao, Jingyin Zhu, Jifan Zhou, Mowei Shen, Tao Gao 0004 |
CogSci | 2 |
| 2022 | Exploring an Imagined "We" in Human Collective Hunting: Joint Commitment within Shared Intentionality
Siyi Gong, Minglu Zhao, Chenya Gu, Jifan Zhou, Mowei Shen, Tao Gao 0004 |
CogSci | 3 |
| 2021 | Modeling Communication to Coordinate Perspectives in Cooperation
Stephanie Stacy, Chenfei Li, Minglu Zhao, Yiling Yun, Qingyi Zhao, Max Kleiman-Weiner, Tao Gao 0004 |
CogSci | 3 |
| 2021 | Sharing is not Needed: Modeling Animal Coordinated Hunting with Reinforcement Learning
Minglu Zhao, Annya L. Dahmani, Ross Richard Perry, Yixin Zhu 0001, Federico Rossano, Tao Gao 0004 |
CogSci | 1 |
| 2020 | Failure Prediction in Datacenters Using Unsupervised Multimodal Anomaly DetectionabstractPredicting hard drive failures in datacenters can help avoid wasting resources and waiting time for recovery. Anomaly detection from sensing data is commonly used for predicting failures. Usually, conventional threshold-based anomaly detection methods consider each sensor independently. However, deciding an optimal threshold for each type of sensors is not trivial, especially for large-scale systems in datacenters. To detect failures that cannot conventionally be detected, multimodal anomaly detection becomes crucial integrating sensing data from different types of sensors. This work proposes a correlation-based multimodal anomaly detection approach. This approach is applied to a Network-Attached Storage (NAS) system with multiple hard disk drives (HDDs) and three sensors, which are a thermal camera, a microphone, and system performance logs. The unimodal results show that the auditory and system performance model can detect temporal anomalies, and the thermal model can detect spatial anomalies. The multimodal results show that even with a simple filter and detection algorithms, the multimodal approach was able to detect failure signs before the real failure and also earlier than the auditory unimodal approach. Minglu Zhao, Reo Furuhata, Mulya Agung, Hiroyuki Takizawa, Tomoya Soma |
IEEE BigData | 1 |
| 2020 | Intuitive Signaling Through an "Imagined We'"
Stephanie Stacy, Qingyi Zhao, Minglu Zhao, Max Kleiman-Weiner, Tao Gao 0004 |
CogSci | 3 |
| 2020 | Bootstrapping an Imagined We for Cooperation
Stephanie Stacy, Minglu Zhao, Gabriel Marquez, Tao Gao 0004 |
CogSci | 3 |