VLDB 2026 Research / reviewers in the wild / expert
Mingfei Chen
dblp:193/7282
· DBLP profile ↗
18ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
5 papers |
Audio and music processing · 51% Rendering · 22% Virtual and augmented reality · 16% | |
| Artificial intelligence
3 papers |
Vision and language · 33% Knowledge representation and reasoning · 21% 3D vision · 20% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
3d spatial reasoning |
0.9 | 1 | 2025 | SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
audio-visual large language model |
0.9 | 1 | 2025 | SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing · NeurIPS 2025 |
Audio and music processing › acoustic rendering
novel-view acoustic synthesis |
0.9 | 1 | 2025 | SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding · CVPR 2025 |
Audio and music processing › spatial audio
spatial audio generation |
0.9 | 1 | 2025 | SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding · CVPR 2025 |
Rendering
neural rendering |
0.8 | 2 | 2024 | Geometry-Guided Progressive NeRF for Generalizable and Efficient Neural Human Rendering · ECCV (23) 2022 AV-Cloud: Spatial Audio Rendering Through Audio-Visual Cloud Splatting · NeurIPS 2024 |
Audio and music processing › spatial audio
sound field reproduction |
0.8 | 1 | 2024 | AV-Cloud: Spatial Audio Rendering Through Audio-Visual Cloud Splatting · NeurIPS 2024 |
Virtual and augmented reality › immersive audio
immersive audio rendering |
0.7 | 1 | 2023 | Be Everywhere - Hear Everything (BEE): Audio Scene Reconstruction by Sparse Audio-Visual Samples · ICCV 2023 |
Computer vision › 3D vision
implicit neural representation |
0.6 | 1 | 2022 | INRAS: Implicit Neural Representation for Audio Scenes · NeurIPS 2022 |
Rendering
human performance rendering |
0.6 | 1 | 2022 | Geometry-Guided Progressive NeRF for Generalizable and Efficient Neural Human Rendering · ECCV (23) 2022 |
Rendering
neural radiance fields |
0.6 | 1 | 2022 | Geometry-Guided Progressive NeRF for Generalizable and Efficient Neural Human Rendering · ECCV (23) 2022 |
Audio and music processing › acoustic simulation
room impulse response rendering |
0.6 | 1 | 2022 | INRAS: Implicit Neural Representation for Audio Scenes · NeurIPS 2022 |
Audio and music processing
spatial audio |
0.6 | 1 | 2022 | INRAS: Implicit Neural Representation for Audio Scenes · NeurIPS 2022 |
Computer vision › Image recognition and object detection
human-object interaction detection |
0.5 | 1 | 2021 | Reformulating HOI Detection As Adaptive Set Prediction · CVPR 2021 |
Computer vision › Vision and language
visual relationship detection |
0.5 | 1 | 2021 | Reformulating HOI Detection As Adaptive Set Prediction · CVPR 2021 |
Computer vision › 3D vision
3d scene understanding |
0.3 | 1 | 2025 | SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing · NeurIPS 2025 |
Robotics › Robot navigation and mapping › robot mapping
global map building |
0.3 | 1 | 2025 | SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing · NeurIPS 2025 |
Computer vision › Video understanding and tracking
object tracking |
0.3 | 1 | 2025 | SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing · NeurIPS 2025 |
Audio and music processing
acoustic rendering |
0.2 | 1 | 2022 | INRAS: Implicit Neural Representation for Audio Scenes · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
implicit neural representation · 1.1panoramic RGB-D · 0.9egocentric spatial track estimation · 0.9cross-modal embedding · 0.9coordinate transformation · 0.9acoustic transfer function learning · 0.9transfer function · 0.8neural splatting · 0.8anchor points · 0.8joint audio-visual representation learning · 0.7neural radiance field · 0.6multicondition training · 0.6geometry-guided progressive learning · 0.6transformer · 0.5set prediction · 0.5attention · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CDMVC: Consensus-Driven Multiview Clustering via Adaptive Hierarchical Matrix Factorization With Dynamic View FusionabstractMulti-view clustering is important for discovering shared patterns from heterogeneous data, but many existing methods are constrained by fixed architectures, rigid fusion strategies, and weak theoretical support. We propose a new framework CDMVC, which integrates adaptive hierarchical matrix factorization with dynamic view fusion. First, an AHMD module is designed to automatically determine the decomposition depth of each view based on its intrinsic complexity, thus avoiding the limitations of manually specified or fixed-depth structures. Second, a consensus-driven fusion mechanism jointly preserves view-specific information and cross-view consistency through an adaptive optimization process with theoretical guarantees. Third, a dynamic weighting strategy updates the contribution of each view according to its structural quality and clustering relevance. Theoretical analysis proves the convergence of CDMVC, and experiments on benchmark datasets demonstrate that CDMVC consistently achieves superior clustering performance while maintaining computational efficiency. Mingfei Chen |
IEEE Internet Things J. | 2 |
| 2026 | MSADroid: A pre-trained Mamba-sparse self-attention model for android malware detection on long system call sequences
Guojun Wang 0001, Mingfei Chen, Yuheng Zhang 0001, Wanyi Gu, Bao Wu |
Inf. Sci. | 3 |
| 2026 | DroidFlow: Learning behavioral densities from Android API sequences for out-of-distribution detection
Wanyi Gu, Guojun Wang 0001, Mingfei Chen, Zhuoyi Wu, Yuheng Zhang 0001 |
J. Syst. Archit. | 3 |
| 2026 | Distributed Strategy Seeking for Aggregative Games Over Digraphs Based on the Out-Degree InformationabstractIn this article, we investigate aggregative games with coupling constraints and local feasibility constraints over digraphs. It is noted that the imbalance introduced by the digraph leads to an inaccurate estimation of the global aggregation function and increases the difficulty of handling coupling constraints. To overcome this issue, a consensus dynamics is designed to estimate the right eigenvector associated with the zero eigenvalue of the Laplacian matrix constructed using the nodes’ out-degree information. By normalizing with the estimated right eigenvector, the imbalance is eliminated, enabling accurate estimation of the global aggregation function. Based on this, a distributed projection-based algorithm is developed, and through the aid of the proposed consensus dynamics, the asymptotic convergence to the Nash equilibrium (NE) is rigorously proven via singular perturbation theory. In particular, when either local feasibility constraints or coupling constraints are absent, the proposed algorithm achieves exponential convergence to the NE. Finally, the effectiveness of the proposed algorithms is validated through simulations on the location problem and Nash–Cournot games. Mingfei Chen, Shuai Liu 0014, Dong Wang 0003, Xian-Ming Zhang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2025 | SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic BindingabstractWe introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the underlying acoustic transfer function that relates the signals acquired at the distributed microphones to the signal at the target viewpoint, using a limited number of known recordings. Unlike existing works, our method does not require constraints or prior knowledge of sound source details. Moreover, our method efficiently adapts to diverse room layouts, reference microphone configurations and unseen environments. To enable this, we introduce a visual-acoustic binding module that learns visual embeddings linked with local acoustic properties from panoramic RGB and depth data. We first leverage these embeddings to optimize the placement of reference microphones in any given scene. During synthesis, we leverage multiple embeddings extracted from reference locations to get adaptive weights for their contribution, conditioned on target viewpoint. We benchmark the task on both publicly available data and real-world settings. We demonstrate significant improvements over existing methods. Mingfei Chen, Israel D. Gebru, Ishwarya Ananthabhotla, Christian Richardt, Dejan Markovic, Jake Sandakly, Steven Krenn, Todd Keebler, Eli Shlizerman, Alexander Richard |
CVPR | 1 |
| 2025 | SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearingabstract3D spatial reasoning in dynamic, audio-visual environments is a cornerstone of human cognition yet remains largely unexplored by existing Audio-Visual Large Language Models (AV-LLMs) and benchmarks, which predominantly focus on static or 2D scenes. We introduce SAVVY-Bench, the first benchmark for 3D spatial reasoning in dynamic scenes with synchronized spatial audio. SAVVY-Bench is comprised of thousands of carefully curated question–answer pairs probing both directional and distance relationships involving static and moving objects, and requires fine-grained temporal grounding, consistent 3D localization, and multi-modal annotation. To tackle this challenge, we propose SAVVY, a novel training-free reasoning pipeline that consists of two stages: (i) Egocentric Spatial Tracks Estimation, which leverages AV-LLMs as well as other audio-visual methods to track the trajectories of key objects related to the query using both visual and spatial audio cues, and (ii) Dynamic Global Map Construction, which aggregates multi-modal queried object trajectories and converts them into a unified global dynamic map. Using the constructed map, a final QA answer is obtained through a coordinate transformation that aligns the global map with the queried viewpoint. Empirical evaluation demonstrates that SAVVY substantially enhances performance of state-of-the-art AV-LLMs, setting a new standard and stage for approaching dynamic 3D spatial reasoning in AV-LLMs. Mingfei Chen, Zijun Cui, Xiulong Liu 0002, Jinlin Xiang, Eli Shlizerman |
NeurIPS | 1 |
| 2024 | A Vulnerability Detection Method for Smart Contract Using Opcode Sequences with Variable Length
Xuelei Liu, Guojun Wang 0001, Mingfei Chen, Peiqiang Li, Jinyao Zhu |
ICIC (8) | 3 |
| 2024 | AV-Cloud: Spatial Audio Rendering Through Audio-Visual Cloud SplattingabstractWe propose a novel approach for rendering high-quality spatial audio for 3D scenes that is in synchrony with the visual stream but does not rely or explicitly conditioned on the visual rendering. We demonstrate that such an approach enables the experience of immersive virtual tourism - performing a real-time dynamic navigation within the scene, experiencing both audio and visual content. Current audio-visual rendering approaches typically rely on visual cues, such as images, and thus visual artifacts could cause inconsistency in the audio quality. Furthermore, when such approaches are incorporated with visual rendering, audio generation at each viewpoint occurs after the rendering of the image of the viewpoint and thus could lead to audio lag that affects the integration of audio and visual streams. Our proposed approach, AV-Cloud, overcomes these challenges by learning the representation of the audio-visual scene based on a set of sparse AV anchor points, that constitute the Audio-Visual Cloud, and are derived from the camera calibration. The Audio-Visual Cloud serves as an audio-visual representation from which the generation of spatial audio for arbitrary listener location can be generated. In particular, we propose a novel module Audio-Visual Cloud Splatting which decodes AV anchor points into a spatial audio transfer function for the arbitrary viewpoint of the target listener. This function, applied through the Spatial Audio Render Head module, transforms monaural input into viewpoint-specific spatial audio. As a result, AV-Cloud efficiently renders the spatial audio aligned with any visual viewpoint and eliminates the need for pre-rendered images. We show that AV-Cloud surpasses current state-of-the-art accuracy on audio reconstruction, perceptive quality, and acoustic effects on two real-world datasets. AV-Cloud also outperforms previous methods when tested on scenes "in the wild". Mingfei Chen, Eli Shlizerman |
NeurIPS | 1 |
| 2024 | TransFront: Bi-path Feature Fusion for Detecting Front-running Attack in Decentralized Finance
Yuheng Zhang 0001, Guojun Wang 0001, Peiqiang Li, Xubin Li, Wanyi Gu, Mingfei Chen, Houji Chen |
TrustCom | 6 |
| 2023 | Be Everywhere - Hear Everything (BEE): Audio Scene Reconstruction by Sparse Audio-Visual SamplesabstractFully immersive and interactive audio-visual scenes are dynamic such that the listeners and the sound emitters move and interact with each other. Reconstruction of an immersive sound experience, as it happens in the scene, requires detailed reconstruction of the audio perceived by the listener at an arbitrary location. The audio at the listener location is a complex outcome of sound propagation through the scene geometry and interacting with surfaces and also the locations of the emitters and the sounds they emit. Due to these aspects, detailed audio reconstruction requires extensive sampling of audio at any potential listener location. This is usually difficult to implement in realistic real-time dynamic scenes. In this work, we propose to circumvent the need for extensive sensors by leveraging audio and visual samples from only a handful of A/V receivers placed in the scene. In particular, we introduce a novel method and end-to-end integrated rendering pipeline which allows the listener to be everywhere and hear everything (BEE) in a dynamic scene in real-time. BEE reconstructs the audio with two main modules, Joint Audio-Visual Representation, and Integrated Rendering Head. The first module extracts the informative audio-visual features of the scene from sparse A/V reference samples, while the second module integrates the audio samples with learned time-frequency transformations to obtain the target sound. Our experiments indicate that BEE outperforms existing methods by a large margin in terms of quality of sound reconstruction, can generalize to scenes not seen in training and runs in real-time speed. Mingfei Chen, Eli Shlizerman |
ICCV | 1 |
| 2023 | Recent progress in image denoising: A training strategy perspectiveabstractAbstract Image denoising is one of the hottest topics in image restoration area, it has achieved great progress both in terms of quantity and quality in recent years, especially after the wide and intensive application of deep neural networks. In many deep learning based image denoising models, the performance can greatly benefit from the prepared clean/noisy image pairs used for model training, however, it also limits the application of these models in real denoising scenes. Therefore, more and more researchers tend to develop models that can be learned without image pairs, namely the denoising models that can be well generalised in real‐world denoising tasks. This motivates to make a survey on the recent development of image denoising methods. In this paper, the typical denoising methods from the perspective of model training are reviewed, the reviewed methods are categorised into four classes: the models need clean/noisy image pairs to train, the models trained on multiple noisy images, the models can be learned from a single noisy image, and the visual transformer based models. The denoising results of different denoisers were compared on some public datasets to discover the performance and advantages. The challenges and future directions in image denoising area are also discussed. Wencong Wu, Mingfei Chen, Yungang Zhang, Yang Yang 0032 |
IET Image Process. | 2 |
| 2022 | Geometry-Guided Progressive NeRF for Generalizable and Efficient Neural Human Rendering
Mingfei Chen, Xiangyu Xu 0002, Yujun Cai, Jiashi Feng, Shuicheng Yan |
ECCV (23) | 1 |
| 2022 | INRAS: Implicit Neural Representation for Audio ScenesabstractThe spatial acoustic information of a scene, i.e., how sounds emitted from a particular location in the scene are perceived in another location, is key for immersive scene modeling. Robust representation of scene's acoustics can be formulated through a continuous field formulation along with impulse responses varied by emitter-listener locations. The impulse responses are then used to render sounds perceived by the listener. While such representation is advantageous, parameterization of impulse responses for generic scenes presents itself as a challenge. Indeed, traditional pre-computation methods have only implemented parameterization at discrete probe points and require large storage, while other existing methods such as geometry-based sound simulations still suffer from inability to simulate all wave-based sound effects. In this work, we introduce a novel neural network for light-weight Implicit Neural Representation for Audio Scenes (INRAS), which can render a high fidelity time-domain impulse responses at any arbitrary emitter-listener positions by learning a continuous implicit function. INRAS disentangles scene’s geometry features with three modules to generate independent features for the emitter, the geometry of the scene, and the listener respectively. These lead to an efficient reuse of scene-dependent features and support effective multi-condition training for multiple scenes. Our experimental results show that INRAS outperforms existing approaches for representation and rendering of sounds for varying emitter-listener locations in all aspects, including the impulse response quality, inference speed, and storage requirements. Mingfei Chen, Eli Shlizerman |
NeurIPS | 2 |
| 2021 | Reformulating HOI Detection As Adaptive Set PredictionabstractDetermining which image regions to concentrate is critical for Human-Object Interaction (HOI) detection. Conventional HOI detectors focus on either detected human and object pairs or pre-defined interaction locations, which limits learning of the effective features. In this paper, we reformulate HOI detection as an adaptive set prediction problem, with this novel formulation, we propose an Adaptive Set-based one-stage framework (AS-Net) with parallel instance and interaction branches. To attain this, we map a trainable interaction query set to an interaction prediction set with transformer. Each query adaptively aggregates the interaction-relevant features from global contexts through multi-head co-attention. Besides, the training process is supervised adaptively by matching each ground-truth with the interaction prediction. Furthermore, we design an effective instance-aware attention module to introduce instructive features from the instance branch into the interaction branch. Our method outperforms previous state-of-the-art methods without any extra human pose and language features on three challenging HOI detection datasets. Especially, we achieve over 31% relative improvement on a large scale HICO-DET dataset. Code is available at https://github.com/yoyomimi/AS-Net. Mingfei Chen, Yue Liao, Si Liu 0001, Zhiyuan Chen 0008, Fei Wang 0032, Chen Qian 0006 |
CVPR | 1 |
| 2019 | Distributed Extremum Seeking for Optimal Resource Allocation and Its Application to Economic Dispatch in Smart GridsabstractThis paper proposes a first-order extremum-seeking algorithm to solve the resource allocation problem, where the specific expression form and gradient information of the local cost functions are not required. Agents take advantage of measurements of local cost functions to minimize the sum of their cost functions while satisfying the resource constraint, where agents exchange the estimated decisions with their neighbors under an undirected and connected graph. Making use of the Lyapunov stability theory and the average analysis method, the convergence of the proposed algorithm to the neighborhood of the optimal solution is presented. In addition, it is obtained that the designed algorithm is semiglobally practically asymptotically stable. Then, the first-order algorithm is extended to the second-order algorithm with low-pass filters, which achieves better convergence performance than the first-order algorithm. Finally, the effectiveness of the proposed algorithm is illustrated by numerical examples and its application to economic dispatch in smart grids. Dong Wang 0003, Mingfei Chen, Wei Wang 0036 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Distributed optimization for multi-agent systems with constraints set and communication time-delay over a directed graph
Dong Wang 0003, Zhu Wang 0010, Mingfei Chen, Wei Wang 0036 |
Inf. Sci. | 3 |
| 2018 | Faulty Line Detection Method Based on Optimized Bistable System for Distribution NetworkabstractThe noneffectively neutral grounded distribution network is called small current to ground system (SCGS) in China. When single-phase to ground fault occurs in SCGS, the fault current is weak, and the noise impairs the feature of fault current, both of which make faulty line detection difficult. This paper presents a faulty line detection method for SCGS, based on optimized bistable system. The proposed method consists of two steps: 1) The optimized bistable system, whose potential function parameters are optimized by particle swarm optimization algorithm, is used to extract transient zero-sequence current (TZSC) in strong noise background; 2) the optimized bistable system and cross correlation coefficient are used to propose a faulty line detection criterion, it based on squared distance, which contains the waveform difference and energy of TZSC. Simulation and field experiments prove that the method can detect faulty line exactly with various fault situations, such as different signal-noise ratios, grounding resistances, initial angles, faulty lines and unbalanced load. Xiaowei Wang 0002, Jie Gao 0005, Mingfei Chen, Xiangxiang Wei, Yanfang Wei, Zhihui Zeng |
IEEE Trans. Ind. Informatics | 3 |
| 2017 | Optimisation of partial collaborative transportation scheduling in supply chain management with 3PL using ACO
Shida Xu, Yanqiu Liu, Mingfei Chen |
Expert Syst. Appl. | 3 |