EDBT 2026 Demo / reviewers in the wild / expert
Wensen Feng
dblp:276/1283
· DBLP profile ↗
17ranked-venue papers
0as first author
16since 2021 · last 2025
0000-0002-7315-337XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 13 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Animus3D: Text-driven 3D Animation via Motion Score DistillationabstractWe present Animus3D , a text-driven 3D animation framework that generates motion field given a static 3D asset and text prompt. Previous methods mostly leverage the vanilla Score Distillation Sampling (SDS) objective to distill motion from pretrained text-to-video diffusion, leading to animations with minimal movement or noticeable jitter. To address this, our approach introduces a novel SDS alternative, Motion Score Distillation (MSD). Specifically, we introduce a LoRA-enhanced video diffusion model that defines a static source distribution rather than pure noise as in SDS, while another inversion-based noise estimation technique ensures appearance preservation when guiding motion. To further improve motion fidelity, we incorporate explicit temporal and spatial regularization terms that mitigate geometric distortions across time and space. Additionally, we propose a motion refinement module to upscale the temporal resolution and enhance fine-grained details, overcoming the fixed-resolution constraints of the underlying video model. Extensive experiments demonstrate that Animus3D successfully animates static 3D assets from diverse text prompts, generating significantly more substantial and detailed motion than state-of-the-art baselines while maintaining high visual integrity. Code will be released upon acceptance. Qi Sun 0005, Can Wang 0007, Jiaxiang Shang, Wensen Feng, Jing Liao 0001 |
SIGGRAPH Asia | 4 |
| 2025 | MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer SynthesisabstractRecent breakthroughs in video generation, powered by large-scale datasets and diffusion techniques, have shown that video diffusion models can function as implicit 4D novel view synthesizers. Nevertheless, current methods primarily concentrate on redirecting camera trajectory within the front view while struggling to generate 360-degree viewpoint changes. In this paper, we focus on human-centric subdomain and present MV-Performer, an innovative framework for creating synchronized novel view videos from monocular full-body captures. To achieve a 360-degree synthesis, we extensively leverage the MVHumanNet dataset and incorporate an informative condition signal. Specifically, we use the camera-dependent normal maps rendered from oriented partial point clouds, which effectively alleviate the ambiguity between seen and unseen observations. To maintain synchronization in the generated videos, we propose a multi-view human-centric video diffusion model that fuses information from the reference video, partial rendering, and different viewpoints. Additionally, we provide a robust inference procedure for in-the-wild video cases, which greatly mitigates the artifacts induced by imperfect monocular depth estimation. Extensive experiments on three datasets demonstrate our MV-Performer’s state-of-the-art effectiveness and robustness, setting a strong model for human-centric 4D novel view synthesis. Code is available at https://github.com/zyhbili/MV-Performer. Yihao Zhi, Chenghong Li, Hongjie Liao, Xihe Yang, Zhengwentai Sun, Xiaodong Cun, Wensen Feng, Xiaoguang Han 0001 |
SIGGRAPH Asia | 8 |
| 2025 | StruGauAvatar: Learning Structured 3D Gaussians for Animatable Avatars From Monocular VideosabstractIn recent years, significant progress has been witnessed in the field of neural 3D avatar reconstruction. Among all related tasks, building an animatable avatar from monocular videos is one of the most challenging ones, yet it also has a wide range of applications. The "animatable" means that we need to transfer any arbitrary and unseen poses onto the avatar and generate new 3D videos. Thanks to the rise of the powerful representation of NeRF, generating a high-fidelity animatable avatar from videos has become easier and more accessible. Despite their impressive visual results, the substantial training and rendering overhead dramatically hamper their applications. 3D Gaussian Splatting, as a timely new representation, has demonstrated its high-quality and high-efficiency rendering. This has led to many concurrent works to introduce 3D-GS to animatable avatar building. Although they demonstrate very high-fidelity renderings for poses similar to the training video frames, poor results are produced when the poses are far from training. We argue that this is primarily because the Gaussian points lack structures. Thus, we suggest involving DMTet to represent the coarse geometry of the avatar. In our representation, the majority of Gaussian points are bound to the mesh vertices, while some free Gaussian is allowed to expand to better fit the given video. Furthermore, we develop a dual-space optimization framework to jointly optimize the DMTet, Gaussian points, and skinning weights under two spaces. In this sense, Gaussian points are transformed in a constrained way, which dramatically improves the generalization ability for unseen poses. This is well demonstrated via extensive experiments. Yihao Zhi, Wanhu Sun, Chongjie Ye, Wensen Feng, Xiaoguang Han 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | GaussReg: Fast 3D Registration with Gaussian Splatting
Yinglin Xu, Yuantao Chen, Wensen Feng, Xiaoguang Han 0001 |
ECCV (15) | 5 |
| 2024 | EyebrowNet: High-Precision Eyebrow Reconstruction and MattingabstractEyebrows play a crucial role in facial reconstruction. However, due to their fine details and the lack of precise eyebrow datasets, traditional methods struggle to extract eyebrow features. To address this, we propose EyebrowNet, a high-precision eyebrow matting network consisting of an optimization phase and an inference phase, to extract high-quality eyebrow masks from blurry inputs. In the optimization phase, we develop a diffusion-based super-resolution network to enhance low-quality eyebrow images. To overcome the limited real eyebrow data, we employ a progressive fine-tuning strategy. In the inference phase, we devise a matting network based on conditional generative adversarial networks. To capture the rich details in real eyebrow images, we employ a synthetic training strategy. Comprehensive experiments prove that our EyebrowNet achieves state-of-the-art results. All source codes and datasets will be released to the public. Wensen Feng, Haoqian Wang |
ICME | 2 |
| 2024 | Learning to match features with discriminative sparse graph neural network
Junxiong Cai, Mingyu Fan, Wensen Feng |
Pattern Recognit. | 4 |
| 2023 | Volumetric 3D Reconstruction with Window-Wise Global Feature AggregationabstractVolumetric 3D reconstruction methods have shown great performance in reconstructing indoor scenarios from monocular videos. However, as such approaches utilize discrete feature voxels to encode the observed scenes, the global feature interaction within and across different voxels is ignored, leading to imperfect reconstructions. To solve this problem, we propose a novel volumetric 3D reconstruction method named VolGARecon. The core portion of VolGARecon includes two parts: first, we use an MLP-based weighted fusion module (WFM) to unproject the extracted features to each voxel, which considers the visibility and is capable to reduce the noise caused by occlusion; second, a 3D transformer module (3DTR) is used to perform window-wise global feature interaction in a local sliding window, which strengthens the feature expression in 3D space and benefits estimating more complete and spatially coherent 3D models. In addition, we propose a multi-dimensional hybrid loss (MHL) that incorporates the 3D supervision in classical volumetric methods and the 2D supervision in novel view synthesis works. Extensive experiments show our method achieves superior performance on multiple datasets. Shihao Ren, Yikang Ding, Jinli Liao, Xinghui Li, Wensen Feng, Xueqian Wang 0001 |
ICASSP | 6 |
| 2023 | Edge-aware Neural Implicit Surface ReconstructionabstractRecently, neural implicit 3D reconstruction in indoor scenarios has achieved impressive performance. Utilizing the volume rendering method and neural implicit representation to learn 3D scenes, such per-scene optimization methods could reconstruct pretty complete models but also suffer from missing details and overly-smoothed reconstructions. In this paper, we propose a novel edge-aware neural implicit surface reconstruction method, named Ea-NeuS, to learn high-quality 3D models with fine details. Specifically, we use the edge of objects to locate the important areas, and propose a simple yet effective edge-guided ray-sampling strategy to learn the 3D models. The aforementioned edge information further guides the normal prior supervision, which helps reduce inaccurate optimization in detailed regions. We additionally use the visibility-aware sparse points to pilot the 3D points sampling along the rays and perform explicit supervision. As a result, our method achieves superior performance compared with existing methods on various scenes. Xinghui Li, Yikang Ding, Xiansong Lai, Shihao Ren, Wensen Feng, Long Zeng 0001 |
ICME | 6 |
| 2023 | SPS: Accurate and Real-Time Semantic Positioning System Based on Low-Cost DEM MapsabstractThis paper presents a Semantic Positioning System (SPS) to enhance the accuracy of mobile device geo-localization in outdoor urban environments. Although the traditional Global Positioning System (GPS) can offer a rough localization, it lacks the necessary accuracy for applications such as Augmented Reality (AR). Our SPS integrates Geographic Information System (GIS) data, GPS signals, and visual image information to estimate the 6 Degree-of-Freedom (DoF) pose through cross-view semantic matching. This approach has excellent scalability to support GIS context with Levels of Detail (LOD). The map data representation is Digital Elevation Model (DEM), a cost-effective aerial map that allows for fast deployment for large-scale areas. However, the DEM lacks geometric and texture details, making it challenging for traditional visual feature extraction to establish pixel/voxel level cross-view correspondences. To address this, we sample observation pixels from the query ground-view image using predicted semantic labels. We then propose an iterative homography estimation method with semantic correspondences. To improve the efficiency of the overall system, we further employ a heuristic search to speedup the matching process. The proposed method is robust, real-time, and automatic. Quantitative experiments on the challenging Bund dataset show that we achieve a positioning accuracy of 73.24%, surpassing the baseline skyline-based method by 20%. Compared with the state-of-the-art semantic-based approach on the Kitti dataset, we improve the positioning accuracy by an average of 5%. Junxiong Cai, Wensen Feng, Haoxiang Chen 0004, Tai-Jiang Mu |
IEEE Trans. Image Process. | 2 |
| 2023 | Seeing Through Darkness: Visual Localization at Night via Weakly Supervised Learning of Domain Invariant FeaturesabstractLong term visual localization has to conquer the problem of matching images with dramatic photometric changes caused by different seasons, natural and man-made illumination changes, etc. Visual localization at night plays a vital role in many applications like autonomous driving and augmented reality, for which extracting keypoints and descriptors with robustness to day-night illumination changes has became the bottleneck. This paper proposes an adversarial learning based solution to harvest from the weakly domain labels of day and night images, along with the point level correspondences among day time images, to achieve robust local feature extraction and description across day-night images. The key idea is to learn a discriminator to distinguish whether a feature map is generated from the day or night images, and simultaneously to adjust the parameters of feature extraction network so as to fool the discriminator. After adversarial training of the discriminator and feature extraction network, the feature extraction network finally reaches a stable status so that the extracted feature maps are robust to day-night photometric changes, based on which day-night domain invariant keypoints and descriptors can be extracted. Compared to existing local feature learning methods, it only requires an additional set of easily captured night images to improve the domain invariance of learned features. Experiments on two challenging benchmarks show the effectiveness of proposed method. In addition, this paper revisits the widely used image matching metrics on HPatches and finds that recall of different methods is highly related to their relative localization performance. Bin Fan 0001, Yuzhu Yang, Wensen Feng, Fuchao Wu, Jiwen Lu, Hongmin Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Deep Blind Image Quality Assessment Powered by Online Hard Example MiningabstractRecently, blind image quality assessment (BIQA) models based on deep neural networks (DNNs) have achieved impressive performance on existing datasets. However, due to the intrinsic imbalance property of the training set, not all distortions or images are handled equally well. Online hard example mining (OHEM) is a promising way to alleviate this issue. Inspired by the recent finding that network pruning disproportionately hampers the model's memorization of a tractable subset, atypical, low-quality, long-tailed samples, that are hard-to-memorize during training and easily “forgotten” during pruning, we propose an effective “plug-and-play” OHEM pipeline, especially for generalizable deep BIQA. Specifically, we train two parallel weight-sharing branches simultaneously, where one is full model and other is a “self-competitor” generated from the full model online by network pruning. Then, we leverage the prediction disagreement between the full model and its pruned variant (i.e., the self-competitor) to expose easily “forgettable” samples, which are therefore regarded as the hard ones. We then enforce the prediction consistency between the full model and its pruned variant to implicitly put more focus on these hard samples, which benefits the full model to recover forgettable information introduced by pruning. Extensive experiments across multiple datasets and BIQA models demonstrate that the proposed OHEM can improve the model performance and generalizability as measured by correlation numbers and group maximum differentiation (gMAD) competition. Our code are available at:https://github.com/wangzhihua520/IQA_with_OHEM Zhihua Wang 0002, Qiuping Jiang, Shanshan Zhao 0001, Wensen Feng, Weisi Lin |
IEEE Trans. Multim. | 4 |
| 2022 | Adaptive Range Guided Multi-view Depth Estimation with Normal Ranking Loss
Yikang Ding, Dihe Huang, Kai Zhang 0012, Zhiheng Li 0001, Wensen Feng |
ACCV (1) | 6 |
| 2022 | ClusterGNN: Cluster-based Coarse-to-Fine Graph Neural Network for Efficient Feature MatchingabstractGraph Neural Networks (GNNs) with attention have been successfully applied for learning visual feature matching. However, current methods learn with complete graphs, resulting in a quadratic complexity in the number of features. Motivated by a prior observation that self- and cross- attention matrices converge to a sparse representation, we propose ClusterGNN, an attentional GNN architecture which operates on clusters for learning the feature matching task. Using a progressive clustering module we adaptively divide keypoints into different subgraphs to reduce redundant connectivity, and employ a coarse-to-fine paradigm for mitigating miss-classification within images. Our approach yields a 59.7% reduction in runtime and 58.4% reduction in memory consumption for dense detection, compared to current state-of-the-art GNN-based matching, while achieving a competitive performance on various computer vision tasks. Junxiong Cai, Yoli Shavit, Tai-Jiang Mu, Wensen Feng, Kai Zhang 0012 |
CVPR | 5 |
| 2022 | Adversarial Learning of Hard Positives for Place RecognitionabstractImage retrieval methods for place recognition learn global image descriptors that are used for fetching geo-tagged images at inference time. Recent works have suggested employing weak and self-supervision for mining hard positives and hard negatives in order to improve localization accuracy and robustness to visibility changes (e.g. in illumination or view point). However, generating hard positives, which is essential for obtaining robustness, is still limited to hard-coded or global augmentations. In this work we propose an adversarial method to guide the creation of hard positives for training image retrieval networks. Our method learns local and global augmentation policies which will increase the training loss, while the image retrieval network is forced to learn more powerful features for discriminating increasingly difficult examples. This approach allows the image retrieval network to generalize beyond the hard examples presented in the data and learn features that are robust to a wide range of variations. Our method achieves state-of-the-art recalls on the Pitts250 and Tokyo 24/7 benchmarks and outperforms recent image retrieval methods on the rOxford and rParis datasets by a noticeable margin. Kai Zhang 0012, Yoli Shavit, Wensen Feng |
IJCNN | 4 |
| 2022 | WT-MVSNet: Window-based Transformers for Multi-view StereoabstractRecently, Transformers have been shown to enhance the performance of multi-view stereo by enabling long-range feature interaction. In this work, we propose Window-based Transformers (WT) for local feature matching and global feature aggregation in multi-view stereo. We introduce a Window-based Epipolar Transformer (WET) which reduces matching redundancy by using epipolar constraints. Since point-to-line matching is sensitive to erroneous camera pose and calibration, we match windows near the epipolar lines. A second Shifted WT is employed for aggregating global information within cost volume. We present a novel Cost Transformer (CT) to replace 3D convolutions for cost volume regularization. In order to better constrain the estimated depth maps from multiple views, we further design a novel geometric consistency loss (Geo Loss) which punishes unreliable areas where multi-view consistency is not satisfied. Our WT multi-view stereo method (WT-MVSNet) achieves state-of-the-art performance across multiple datasets and ranks $1^{st}$ on Tanks and Temples benchmark. Code will be available upon acceptance. Jinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang, Shihao Ren, Wensen Feng, Kai Zhang 0012 |
NeurIPS | 7 |
| 2022 | Learning Semantic-Aware Local Features for Long Term Visual LocalizationabstractExtracting robust and discriminative local features from images plays a vital role for long term visual localization, whose challenges are mainly caused by the severe appearance differences between matching images due to the day-night illuminations, seasonal changes, and human activities. Existing solutions resort to jointly learning both keypoints and their descriptors in an end-to-end manner, leveraged on large number of annotations of point correspondence which are harvested from the structure from motion and depth estimation algorithms. While these methods show improved performance over non-deep methods or those two-stage deep methods, i.e., detection and then description, they are still struggled to conquer the problems encountered in long term visual localization. Since the intrinsic semantics are invariant to the local appearance changes, this paper proposes to learn semantic-aware local features in order to improve robustness of local feature matching for long term localization. Based on a state of the art CNN architecture for local feature learning, i.e., ASLFeat, this paper leverages on the semantic information from an off-the-shelf semantic segmentation network to learn semantic-aware feature maps. The learned correspondence-aware feature descriptors and semantic features are then merged to form the final feature descriptors, for which the improved feature matching ability has been observed in experiments. In addition, the learned semantics embedded in the features can be further used to filter out noisy keypoints, leading to additional accuracy improvement and faster matching speed. Experiments on two popular long term visual localization benchmarks (Aachen Day and Night v1.1, Robotcar Seasons) and one challenging indoor benchmark (InLoc) demonstrate encouraging improvements of the localization accuracy over its counterpart and other competitive methods. Bin Fan 0001, Wensen Feng, Huayan Pu, Yuzhu Yang, Qingqun Kong, Fuchao Wu, Hongmin Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Controllable Continuous Gaze RedirectionabstractIn this work, we present interpGaze, a novel framework for controllable gaze redirection that achieves both precise redirection and continuous interpolation. Given two gaze images with different attributes, our goal is to redirect the eye gaze of one person into any gaze direction depicted in the reference image or to generate continuous intermediate results. To accomplish this, we design a model including three cooperative components: an encoder, a controller and a decoder. The encoder maps images into a well-disentangled and hierarchically-organized latent space. The controller adjusts the magnitudes of latent vectors to the desired strength of corresponding attributes by altering a control vector. The decoder converts the desired representations from the attribute space to the image space. To facilitate covering the full space of gaze directions, we introduce a high-quality gaze image dataset with a large range of directions, which also benefits researchers in related areas. Extensive experimental validation and comparisons to several baseline methods show that the proposed interpGaze outperforms state-of-the-art methods in terms of image quality and redirection precision. Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue, Wensen Feng |
ACM Multimedia | 4 |