EDBT 2026 Demo / reviewers in the wild / expert
Ningning Ma
dblp:30/2310
· DBLP profile ↗
16ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D TrajectoryabstractData scarcity continues to be a critical bottleneck in the field of robotic manipulation, limiting the ability to train robust and generalizable models. While diffusion models provide a promising approach to synthesizing realistic robotic manipulation videos, their effectiveness hinges on the availability of precise and reasonable control instructions. Current methods primarily rely on 2D trajectories as instruction prompts, which inherently face issues with 3D spatial ambiguity. In this work, we present a novel framework named ManipDreamer3Dfor generating plausible 3D-aware robotic manipulation videos from the input image and the text instruction. Our method combines 3D trajectory planning with a reconstructed 3D occupancy map created from a third-person perspective, along with a novel trajectory-to-video diffusion model. Specifically, ManipDreamer3D first reconstructs the 3D occupancy representation from the input image and then computes an optimized 3D end-effector trajectory, minimizing path length, avoiding collisions and retiming. Next, we employ a latent editing technique to create video sequences from the initial image latent, text instruction and the optimized 3D trajectory. This process conditions our specially trained trajectory-to-video diffusion model to produce robotic pick-and-place videos. Our method significantly reduces human intervention requirements by autonomously planing plausible 3D trajectories. Experimental results demonstrate its superior visual quality and precision. Ying Li 0128, Xiaobao Wei, Xiaowei Chi, Zhongyu Zhao, Hao Wang 0073, Ningning Ma, Ming Lu 0002, Sirui Han |
AAAI | 7 |
| 2026 | SparseStreet: Sparse Gaussian Splatting for Real-Time Street Scene SimulationabstractWhile 3D Gaussian Splatting has shown promising results in street scene reconstruction, existing methods require massive numbers of Gaussian primitives to capture fine details, leading to prohibitive storage costs and slow rendering speeds. We observe that dynamic objects (e.g., vehicles and pedestrians) demand high-fidelity representations to maintain temporal consistency, while static background regions often contain substantial redundancy. Motivated by this, we propose SparseStreet, a general compression framework specifically designed for street scenes. First, we introduce a node-based learnable pruning strategy that systematically removes low-contributing Gaussian primitives while preserving visually critical regions. Second, after the scene representation stabilizes, we apply background compression, further reducing redundancy in static regions. Our method effectively preserves the geometry and appearance of dynamic objects while significantly reducing the total number of Gaussian primitives. Extensive experiments on the Waymo and nuScenes demonstrate that SparseStreet achieves up to 80% compression ratio with minimal quality degradation, enabling resource-efficient, high-fidelity dynamic scene reconstruction. Project website: https://sparsestreet.github.io/. Qingpo Wuwu, Xiaobao Wei, Peng Chen 0046, Zhongyu Zhao, Hao Wang 0073, Ming Lu 0002, Ningning Ma, Shanghang Zhang |
ICMR | 8 |
| 2025 | MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual EncodersabstractVisual encoders are fundamental components in vision-language models (VLMs), each showcasing unique strengths derived from various pre-trained visual foundation models. To leverage the various capabilities of these encoders, recent studies incorporate multiple encoders within a single VLM, leading to a considerable increase in computational cost. In this paper, we present Mixture-of-Visual-Encoder Knowledge Distillation (MoVEKD), a novel framework that distills the unique proficiencies of multiple vision encoders into a single, efficient encoder model. Specifically, to mitigate conflicts and retain the unique characteristics of each teacher encoder, we employ low-rank adaptation (LoRA) and mixture-of-experts (MoEs) to selectively activate specialized knowledge based on input features, enhancing both adaptability and efficiency. To regularize the KD process and enhance performance, we propose an attention-based distillation strategy that adaptively weighs the different encoders and emphasizes valuable visual tokens, reducing the burden of replicating comprehensive but distinct features from multiple teachers. Comprehensive experiments on popular VLMs, such as LLaVA and LLaVA-NeXT, validate the effectiveness of our method. Our code is available at: https://github.com/hey-cjj/MoVE-KD. Jiajun Cao, Yuan Zhang 0020, Tao Huang 0020, Ming Lu 0002, Qizhe Zhang, Ruichuan An, Ningning Ma, Shanghang Zhang |
CVPR | 7 |
| 2025 | EMD: Explicit Motion Modeling for High-Quality Street Gaussian SplattingabstractPhotorealistic reconstruction of street scenes is essential for developing real-world simulators in autonomous driving. While recent methods based on 3D/4D Gaussian Splatting (GS) have demonstrated promising results, they still encounter challenges in complex street scenes due to the unpredictable motion of dynamic objects. Current methods typically decompose street scenes into static and dynamic objects, learning the Gaussians in either a supervised manner (e.g., w/ 3D bounding-box) or a self-supervised manner (e.g., w/o 3D bounding-box). However, these approaches do not effectively model the motions of dynamic objects (e.g., the motion speed of pedestrians is clearly different from that of vehicles), resulting in suboptimal scene decomposition. To address this, we propose Explicit Motion Decomposition (EMD), which models the motions of dynamic objects by introducing learnable motion embeddings to the Gaussians, enhancing the decomposition in street scenes. The proposed plug-and-play EMD module compensates for the lack of motion modeling in self-supervised street Gaussian splatting methods. We also introduce tailored training strategies to extend EMD to supervised approaches. Comprehensive experiments demonstrate the effectiveness of our method, achieving state-of-the-art novel view synthesis performance in self-supervised settings. The code is available at: https://qingpowuwu.github.io/emd. Xiaobao Wei, Qingpo Wuwu, Zhongyu Zhao, Zhuangzhe Wu, Ming Lu 0002, Ningning Ma, Shanghang Zhang |
ICCV | 7 |
| 2025 | Boosting 3D Object Detection via Self-Distilling Introspective Dataabstract3D object detection is a fundamental yet critical task for autonomous driving. In this paper, we investigate a novel self-distilling paradigm by proposing Self-distilling Introspective Data (SID) to boost the accuracy of 3D object detection in both LiDAR-based and LiDAR-Camera-based scenarios. The proposed SID significantly improves the applicability of the distillation approach since it does not require extra training data or complex teacher network design. Specifically, we first employ an introspective data augmentation method to enrich object-aware information in sparse point clouds through geometric or semantic injection. We then utilize this enhanced data to train a robust teacher model. In contrast to traditional distillation that relies on larger models to enhance the representations of smaller ones, the teacher model in SID shares the same architecture as the student model but exhibits exceptionally high discriminative ability. This enables the effective transfer of rich feature representations to the student model. Rooted on such a scheme, when conducting LiDAR-based detectors, SID significantly enhances the semantic representation capabilities of sparse point clouds. Additionally, in the LiDAR-Camera-based setting, SID also effectively supervises the fusion of the two modalities at the feature level, ensuring more reasonable cross-modal learning. Extensive experiments show the proposed SID improves a variety of detectors. For the LiDAR-based detector, the SID gains 2.31% mAP improvements for the hard objects in KITTI, while 1.76% NDS improvements on nuScenes. For the LiDAR-Camera-based detectors, the SID boosts the detection accuracy significantly, with 1.5% mAP promotion on KITTI and 2.15% NDS improvements on the nuScenes benchmark. Chaoqun Wang 0012, Yiran Qin, Zijian Kang, Ningning Ma, Yukai Shi, Zhen Li 0026, Ruimao Zhang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | TP-LSM: visual temporal pyramidal time modeling network to multi-label action detection in image-based AIabstractDense multi-label action detection is a challenging task in the field of visual action, where multiple actions occur simultaneously in different time spans, hence accurately assessing the short-term and long-term temporal dependencies between actions is crucial for action detection. There is an urgent need for an effective temporal modeling technology to detect the temporal dependence of actions in videos and efficiently learn long-term and short-term action information. This paper proposes a new method based on temporal pyramid and long short-term time modeling for multi-label action detection, which combines hierarchical structure with pyramid feature hierarchy for dense multi-label temporal action detection. By using the expansion and compression convolution module (SEC) and external attention for time modeling, we focus on the temporal relationships of long and short-term actions at each stage. We then integrate hierarchical pyramid features to achieve accurate detection of actions at different temporal resolution scales. We evaluated the performance of the model on dense multi-label benchmark datasets, and achieved mAP of 47.3% and 36.0% on the MultiTHUMOS and TSU datasets, which outperforms 2.7% and 2.3% on the current state-of-the-art results. The code is available at https://github.com/Yoona6371/TP-LSM . Haojie Gao, Peishun Liu, Zikang Yan, Ningning Ma, Ruichun Tang |
Vis. Comput. | 5 |
| 2024 | Toward Accurate Camera-based 3D Object Detection via Cascade Depth Estimation and CalibrationabstractRecent camera-based 3D object detection is limited by the precision of transforming from image to 3D feature spaces, as well as the accuracy of object localization within the 3D space. This paper aims to address such a fundamental problem of camera-based 3D object detection: How to effectively learn depth information for accurate feature lifting and object localization. Different from previous methods which directly predict depth distributions by using a supervised estimation model, we propose a cascade framework consisting of two depth-aware learning paradigms. First, a depth estimation (DE) scheme leverages relative depth information to realize the effective feature lifting from 2D to 3D spaces. Furthermore, a depth calibration (DC) scheme introduces depth reconstruction to further adjust the 3D object localization perturbation along the depth axis. In practice, the DE is explicitly realized by using both the absolute and relative depth optimization loss to promote the precision of depth prediction, while the capability of DC is implicitly embedded into the detection Transformer through a depth denoising mechanism in the training phase. The entire model training is accomplished through an end-to-end manner. We propose a baseline detector and evaluate the effectiveness of our proposal with +2.2%/+2.7% NDS/mAP improvements on NuScenes benchmark, and gain a comparable performance with 55.9%/45.7% NDS/mAP. Furthermore, we conduct extensive experiments to demonstrate its generality based on various detectors with about +2% NDS improvements. Chaoqun Wang 0012, Yiran Qin, Zijian Kang, Ningning Ma, Ruimao Zhang |
ICRA | 4 |
| 2023 | Lift3D: Synthesize 3D Training Data by Lifting 2D GAN to 3D Generative Radiance FieldabstractThis work explores the use of 3D generative models to synthesize training data for 3D vision tasks. The key requirements of the generative models are that the generated data should be photorealistic to match the real-world scenarios, and the corresponding 3D attributes should be aligned with given sampling labels. However, we find that the recent NeRF-based 3D GANs hardly meet the above requirements due to their designed generation pipeline and the lack of explicit 3D supervision. In this work, we propose Lift3D, an inverted 2D-to-3D generation framework to achieve the data generation objectives. Lift3D has several merits compared to prior methods: (1) Unlike previous 3D GANs that the output resolution is fixed after training, Lift3D can generalize to any camera intrinsic with higher resolution and photorealistic output. (2) By lifting well-disentangled 2D GAN to 3D object NeRF, Lift3D provides explicit 3D information of generated objects, thus offering accurate 3D annotations for downstream tasks. We evaluate the effectiveness of our framework by augmenting autonomous driving datasets. Experimental results demonstrate that our data generation framework can effectively improve the performance of 3D object detectors. Code: len-li.github.io/lift3d-web Leheng Li, Qing Lian, Luozhou Wang, Ningning Ma, Ying-Cong Chen |
CVPR | 4 |
| 2023 | SupFusion: Supervised LiDAR-Camera Fusion for 3D Object DetectionabstractLiDAR-Camera fusion-based 3D detection is a critical task for automatic driving. In recent years, many LiDAR-Camera fusion approaches sprung up and gained promising performances compared with single-modal detectors, but always lack carefully designed and effective supervision for the fusion process. In this paper, we propose a novel training strategy called SupFusion, which provides an auxiliary feature level supervision for effective LiDAR-Camera fusion and significantly boosts detection performance. Our strategy involves a data enhancement method named Polar Sampling, which densifies sparse objects and trains an assistant model to generate high-quality features as the supervision. These features are then used to train the LiDAR-Camera fusion model, where the fusion feature is optimized to simulate the generated high-quality features. Furthermore, we propose a simple yet effective deep fusion module, which contiguously gains superior performance compared with previous fusion methods with SupFusion strategy. In such a manner, our proposal shares the following advantages. Firstly, SupFusion introduces auxiliary feature-level supervision which could boost LiDAR-Camera detection performance without introducing extra inference costs. Secondly, the proposed deep fusion could continuously improve the detector’s abilities. Our proposed SupFusion and deep fusion module is plug-and-play, we make extensive experiments to demon-strate its effectiveness. Specifically, we gain around 2% 3D mAP improvements on KITTI benchmark based on multiple LiDAR-Camera 3D detectors. Our code is available at https://github.com/IranQin/SupFusion. Yiran Qin, Chaoqun Wang 0012, Zijian Kang, Ningning Ma, Zhen Li 0026, Ruimao Zhang |
ICCV | 4 |
| 2021 | RepVGG: Making VGG-Style ConvNets Great AgainabstractWe present a simple but powerful architecture of convolutional neural network, which has a VGG-like inference-time body composed of nothing but a stack of 3 × 3 convolution and ReLU, while the training-time model has a multi-branch topology. Such decoupling of the training-time and inference-time architecture is realized by a structural re-parameterization technique so that the model is named RepVGG. On ImageNet, RepVGG reaches over 80% top-1 accuracy, which is the first time for a plain model, to the best of our knowledge. On NVIDIA 1080Ti GPU, RepVGG models run 83% faster than ResNet-50 or 101% faster than ResNet-101 with higher accuracy and show favorable accuracy-speed trade-off compared to the state-of-the-art models like EfficientNet and RegNet. The code and trained models are available at https://github.com/megvii-model/RepVGG. Xiaohan Ding, Xiangyu Zhang 0005, Ningning Ma, Jungong Han, Guiguang Ding, Jian Sun 0001 |
CVPR | 3 |
| 2021 | Activate or Not: Learning Customized ActivationabstractWe present a simple, effective, and general activation function we term ACON which learns to activate the neurons or not. Interestingly, we find Swish, the recent popular NAS-searched activation, can be interpreted as a smooth approximation to ReLU. Intuitively, in the same way, we approximate the more general Maxout family to our novel ACON family, which remarkably improves the performance and makes Swish a special case of ACON. Next, we present meta-ACON, which explicitly learns to optimize the parameter switching between non-linear (activate) and linear (inactivate) and provides a new design space. By simply changing the activation function, we show its effectiveness on both small models and highly optimized large models (e.g. it improves the ImageNet top-1 accuracy rate by 6.7% and 1.8% on MobileNet-0.25 and ResNet-152, respectively). Moreover, our novel ACON can be naturally transferred to object detection and semantic segmentation, showing that ACON is an effective alternative in a variety of tasks. Code is available at https://github.com/nmaac/acon. Ningning Ma, Xiangyu Zhang 0005, Jian Sun 0001 |
CVPR | 1 |
| 2020 | WeightNet: Revisiting the Design Space of Weight Networks
Ningning Ma, Xiangyu Zhang 0005, Jiawei Huang 0003, Jian Sun 0001 |
ECCV (15) | 1 |
| 2020 | Funnel Activation for Visual Recognition
Ningning Ma, Xiangyu Zhang 0005, Jian Sun 0001 |
ECCV (11) | 1 |
| 2018 | ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design
Ningning Ma, Xiangyu Zhang 0005, Hai-Tao Zheng 0002, Jian Sun 0001 |
ECCV (14) | 1 |
| 2018 | Weakly-supervised image captioning based on rich contextual information
Hai-Tao Zheng 0002, Zhe Wang 0010, Ningning Ma, Jin-Yuan Chen, Xi Xiao 0001, Arun Kumar Sangaiah |
Multim. Tools Appl. | 3 |
| 2015 | Capturing AU-Aware Facial Features and Their Latent Relations for Emotion Recognition in the WildabstractThe Emotion Recognition in the Wild (EmotiW) Challenge has been held for three years. Previous winner teams primarily focus on designing specific deep neural networks or fusing diverse hand-crafted and deep convolutional features. They all neglect to explore the significance of the latent relations among changing features resulted from facial muscle motions. In this paper, we study this recognition challenge from the perspective of analyzing the relations among expression-specific facial features in an explicit manner. Our method has three key components. First, we propose a pair-wise learning strategy to automatically seek a set of facial image patches which are important for discriminating two particular emotion categories. We found these learnt local patches are in part consistent with the locations of expression-specific Action Units (AUs), thus the features extracted from such kind of facial patches are named AU-aware facial features. Second, in each pair-wise task, we use an undirected graph structure, which takes learnt facial patches as individual vertices, to encode feature relations between any two learnt facial patches. Finally, a robust emotion representation is constructed by concatenating all task-specific graph-structured facial feature relations sequentially. Extensive experiments on the EmotiW 2015 Challenge testify the efficacy of the proposed approach. Without using additional data, our final submissions achieved competitive results on both sub-challenges including the image based static facial expression recognition (we got 55.38% recognition accuracy outperforming the baseline 39.13% with a margin of 16.25%) and the audio-video based emotion recognition (we got 53.80% recognition accuracy outperforming the baseline 39.33% and the 2014 winner team's final result 50.37% with the margins of 14.47% and 3.43%, respectively). Anbang Yao, Junchao Shao, Ningning Ma, Yurong Chen 0001 |
ICMI | 3 |