VLDB 2026 Research / reviewers in the wild / expert
Haifeng Wu
dblp:22/3566
· DBLP profile ↗
16ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-authorSystems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | S-INF: Towards Realistic Indoor Scene Synthesis via Scene Implicit Neural FieldabstractLearning-based methods have become increasingly popular in 3D indoor scene synthesis (ISS), showing superior performance over traditional optimization-based approaches. These learning-based methods typically model distributions on simple yet explicit scene representations using generative models. However, due to the oversimplified explicit representations that overlook detailed information and the lack of guidance from multimodal relationships within the scene, most learning-based methods struggle to generate indoor scenes with realistic object arrangements and styles. In this paper, we introduce a new method, Scene Implicit Neural Field (S-INF), for indoor scene synthesis, aiming to learn meaningful representations of multimodal relationships, to enhance the realism of indoor scene synthesis. S-INF assumes that the scene layout is often related to the object-detailed information. It disentangles the multimodal relationships into scene layout relationships and detailed object relationships, fusing them later through implicit neural fields (INFs). By learning specialized scene layout relationships and projecting them into S-INF, we achieve a realistic generation of scene layout. Additionally, S-INF captures dense and detailed object relationships through differentiable rendering, ensuring stylistic consistency across objects. Through extensive experiments on the benchmark 3D-FRONT dataset, we demonstrate that our method consistently achieves state-of-the-art performance under different types of ISS. Zixi Liang, Haifeng Wu, Wen Li 0001, Lixin Duan |
AAAI | 3 |
| 2025 | GeoDepth: From Point-to-Depth to Plane-to-Depth Modeling for Self-Supervised Monocular Depth EstimationabstractSelf-supervised monocular depth estimation has long been treated as a point-wise prediction problem, where the depth of each pixel is usually estimated independently. However, artifacts are often observed in the estimated depth map, e.g., depth values for points located in the same region may jump dramatically. To address this issue, we propose a novel self-supervised monocular depth estimation framework called GeoDepth, where we explore the intrinsic geometric representation in 3D scenes for producing accurate and continuous depth maps. In particular, we model the complex 3D scene as a collection of planes with varying sizes, where each plane is characterized by a unique set of parameters, namely planar normal (indicating plane orientation) and planar offset (defining the perpendicular distance from the camera center to the plane). Under this modeling, points in the same plane are enforced to share a unique representation and their depth variations related only to pixel coordinates, thus this geometric relationship can be exploited to regularize the depth variations of these points. To this end, we design a structured plane generation module that introduces spatio-temporal geometric cues and the plane uniqueness principle to recover the correct scene plane representation. In addition, we develop a depth discontinuity module to identify depth discontinuity regions and subsequently optimize them. Our experiments on the KITTI and NYUv2 datasets demonstrate that GeoDepth achieves state-of-the-art performance, with additional tests on Make3D and ScanNet validating its generalization capabilities. Haifeng Wu, Shuhang Gu, Lixin Duan, Wen Li 0001 |
CVPR | 1 |
| 2025 | P2WNet: Homography Estimation for Part-To-Whole and Cross-Modality ScenariosabstractDeep learning-based homography estimation has achieved remarkable advances in recent years. However, existing methods face limitations in the Part-To-Whole (P2W) scenario, where the template image corresponds to a small portion of the search image, as their designs are tailored to image pairs with similar content and limited displacement. To address this issue, we propose P2WNet, a novel framework for part-to-whole and cross-modality homography estimation. First, we tailor a pseudo-siamese encoder to handle cross-modal inputs and incorporate a transformer-based cascade for feature enhancement. Furthermore, we design a novel P2W matching module to capture and represent the correspondences between image pairs in the Part-To-Whole scenario. These robust features are fed into a prediction module to estimate the homography matrix. Additionally, we propose a dataset to validate our model in P2W scenario. Experiments on both our dataset and public benchmarks (DroneVehicle, GoogleMap, MSCOCO) demonstrate that P2WNet achieves superior performance in the P2W scenario and performs competitively in conventional scenarios. Code is available at https://github.com/xuanxh1/P2WNET. ShangXuan Xie, Haifeng Wu, Wen Li 0001, Lixin Duan |
ICME | 2 |
| 2025 | Dynamic local affine transformation for enhanced text-to-image generation with GANs
Qiang Lan, Haifeng Wu |
Vis. Comput. | 2 |
| 2024 | Fast monte-carlo clustering for signal separation of RFID collision tags at physical layer
Yu Zeng 0002, Hongwei Ding 0005, Haifeng Wu |
Wirel. Networks | 3 |
| 2023 | Weighted hybrid order total variation model using structure tensor for image denoising
Wanru Xu, Haifeng Wu, Ali Abdullah Yahya |
Multim. Tools Appl. | 3 |
| 2023 | Multi-aspect Understanding with Cooperative Graph Attention Networks for Medical Dialogue Information ExtractionabstractMedical dialogue information extraction is an important but challenging task for Electronic Medical Records. Existing medical information extraction methods ignore the crucial information of sentence and multi-level dependency in dialogue, which limits their effectiveness for capturing essential medical information. To address these issues, we present a novel Multi-aspect Understanding with Cooperative Graph Attention Networks for Medical Dialogue Information Extraction to capture multi-aspect sentence information and multi-level dependency information from the dialogue. First, we propose the multi-aspect sentence encoder to capture various features from different perspectives. Second, we propose double graph attention networks to model the dependency features from intra-window and inter-window, respectively. Extensive experiments on a benchmark dataset have well-validated the effectiveness of the proposed method. Haifeng Wu |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2021 | DTMNet: A Discrete Tchebichef Moments-based Deep Neural Network for Multi-focus Image FusionabstractCompared with traditional methods, the deep learning-based multi-focus image fusion methods can effectively improve the performance of image fusion tasks. However, the existing deep learning-based methods encounter a common issue of a large number of parameters, which leads to the deep learning models with high time complexity and low fusion efficiency. To address this issue, we propose a novel discrete Tchebichef moment-based Deep neural network, termed as DTMNet, for multi-focus image fusion. The proposed DTMNet is an end-to-end deep neural network with only one convolutional layer and three fully connected layers. The convolutional layer is fixed with DTM co-efficients (DTMConv) to extract high/low-frequency information without learning parameters effectively. The three fully connected layers have learnable parameters for feature classification. Therefore, the proposed DTMNet for multi-focus image fusion has a small number of parameters (0.01M paras vs. 4.93M paras of regular CNN) and high computational efficiency (0.32s vs. 79.09s by regular CNN to fuse an image). In addition, a large-scale multi-focus image dataset is synthesized for training and verifying the deep learning model. Experimental results on three public datasets demonstrate that the proposed method is competitive with or even outperforms the state-of-the-art multi-focus image fusion methods in terms of subjective visual perception and objective evaluation metrics. Bin Xiao 0002, Haifeng Wu, Xiuli Bi |
ICCV | 2 |
| 2020 | State Estimation of Hemodynamic Model for fMRI Under Confounds: SSM MethodabstractThrough hemodynamic models, the change of neuronal state can be estimated from functional magnetic resonance imaging (fMRI) signals. Usually, there are confounds in the fMRI signal, which will degrade the performance of the estimation for the neuronal state change. For the reason, this paper introduces a state-space model with confounds, from a conventional hemodynamic model. In this model, a successive state estimation method requires a state value vector, an error covariance, an innovation covariance, and a cross covariance to be re-derived. Thus, a confounds square-root cubature Kalman smoothing (CSCKS) algorithm is proposed in this paper. We use a Balloon-Windkessel model to generate simulation data and add confounds signals to evaluate the performance of the proposed algorithm. The experiment results show that when the signal-to-interference ratio is less than 21 dB, the CSCKS proposed in this paper reduced estimation error to 16%, whereas the traditional algorithm reduced it to only 73%. Haifeng Wu, Mingzhi Lu, Yu Zeng 0002 |
IEEE J. Biomed. Health Informatics | 1 |
| 2015 | Capture-Aware Estimation for Large-Scale RFID Tags IdentificationabstractHow to estimate the number of passive radio frequency identification (RFID) tags and the occurrence probability of capture effect is very important for a dynamic frame length Aloha RFID system with capture effect. The estimation would relate to setting an optimal frame length, which makes tag identification achieve higher efficiency. Under large-scale tags identification environment, the number of tags may be much greater than an initial frame length. In this scenario, existing estimates do not work well. In this letter, we propose a novel estimation method for the large-scale tags identification. The proposed method could adjust the initial frame length matched to the number of tags from only the first several slots in the frame. The advantage of the proposed method is to work better even when the number of tags is much greater. Numerical results show that, the proposed method has lower estimation errors under the large-scale tag identification. After setting an optimal frame length from the estimated results of the proposed method, furthermore, we could obtain higher identification efficiency. Yang Wang 0066, Haifeng Wu, Yu Zeng 0002 |
IEEE Signal Process. Lett. | 2 |
| 2013 | Binary Tree Slotted ALOHA for Passive RFID Tag AnticollisionabstractIn order to enhance the efficiency of radio frequency identification (RFID) and lower system computational complexity, this paper proposes three novel tag anticollision protocols for passive RFID systems. The three proposed protocols are based on a binary tree slotted ALOHA (BTSA) algorithm. In BTSA, tags are randomly assigned to slots of a frame and if some tags collide in a slot, the collided tags in the slot will be resolved by binary tree splitting while the other tags in the subsequent slots will wait. The three protocols utilize a dynamic, an adaptive, and a splitting method to adjust the frame length to a value close to the number of tags, respectively. For BTSA, the identification efficiency can achieve an optimal value only when the frame length is close to the number of tags. Therefore, the proposed protocols efficiency is close to the optimal value. The advantages of the protocols are that, they do not need the estimation of the number of tags, and their efficiency is not affected by the variance of the number of tags. Computer simulation results show that splitting BTSA's efficiency can achieve 0.425, and the other two protocols efficiencies are about 0.40. Also, the results show that the protocols efficiency curves are nearly horizontal when the number of tags increases from 20 to 4,000. Haifeng Wu, Yu Zeng 0002, Jihua Feng |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Efficient Framed Slotted Aloha Protocol for RFID Tag AnticollisionabstractIn this paper, we propose a novel efficient frame slotted aloha (EFSA) protocol for radio frequency identification (RFID) tag anticollision in this paper. After successfully identifying each tag, the EFSA protocol will allocate the tag a slot number, which signifies when the tag could be identified during a read cycle. When no tags arrive and leave, idle slots and collision slots will not be produced in subsequent read cycles. In addition, if there is a collision slot in the EFSA protocol, colliding tags in the collision slot will be resolved byQ-ary splitting whereQis equal to the estimated number of colliding tags, while the other unidentified tags is in a waiting state until the colliding tags are successfully resolved. Since allocation of tags to slots is not random in the EFSA protocol, conventional tag quantity estimates are not suitable. Therefore, we also propose a novel tag quantity estimate for the EFSA protocol. Simulation results show that the EFSA protocol outperforms conventional protocols, in term of time slots of reidentifying tags, and the proposed estimate error is less than the conventional estimates in the EFSA protocol. Haifeng Wu, Yu Zeng 0002 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2010 | Bayesian Tag Estimate and Optimal Frame Length for Anti-Collision Aloha RFID SystemabstractIn a dynamic framed slotted aloha RFID system, the key technique can be divided into two parts: precisely estimating tag quantity and determining an optimal frame length. For estimating tag quantity, this paper uses three risk functions to propose three Bayesian estimates, and improves the estimates to reduce computational complexity. For determining an optimal frame length, this paper derives an optimal frame length, which can make the system achieve maximum channel usage efficiency under the condition that the durations of an idle, a collision and a successful slot are not identical. In our simulations, comparison with several conventional tag estimates shows that the proposed Bayesian tag estimates have less error. In addition, the improved estimates have lower computational complexity and their estimate performance is not reduced. The simulations results also indicate that the derived optimal frame length guarantees the maximum channel usage efficiency. Haifeng Wu, Yu Zeng 0002 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2009 | Pareto cooperative coevolutionary genetic algorithm using reference sharing collaborationabstractEpistasis has been a well-known hard problem in optimization solved by evolution, especially by cooperative coevolution. Standard cooperative coevolution usually gets worse performance than standard evolution for optimization problems with epistasis. In this work, we propose a much improved version of cooperative coevolutionary model by using reference sharing collaboration. Pareto dominance is used for measuring the performance of individuals in our algorithm. We evaluate and compare our method with standard evolution and cooperative coevolution on a suite of test problems with and without epistasis interaction. Our experimental results show that the proposed algorithm outperforms the compared methods in most of the cases, and especially, it is superior to the standard evolution to handle epistasis. Haifeng Wu |
GECCO | 2 |
| 2008 | Evolving Efficient Connection for the Design of Artificial Neural Networks
Haifeng Wu |
ICANN (2) | 2 |
| 2008 | Support vector machines for traffic signs recognitionabstractIn many traffic sign recognition system, one of the main tasks is to classify the shapes of traffic sign. In this paper, we have developed a shape-based classification model by using support vector machines. We focused on recognizing seven categories of traffic sign shapes and five categories of speed limit signs. Two kinds of features, binary image and Zernike moments, were used for representing the data to the SVM for training and test. We compared and analyzed the performances of the SVM recognition model using different feature representations and different kernels and SVM types. Our experimental data sets consisted of 350 traffic sign shapes and 250 speed limit signs. Experimental results have shown excellent results, which have achieved 100% accuracy on sign shapes classification and 99% accuracy on speed limit signs classification. The performance of SVM model highly depends on the choice of model parameters. Two search algorithms, grid search and simulated annealing search have been implemented to improve the performances of our classification model. The SVM model were also shown to be more effective than Fuzzy ARTMAP model. Haifeng Wu, Hasan Fleyeh |
IJCNN | 2 |