EDBT 2026 Demo / reviewers in the wild / expert
Hengyi Wang
dblp:215/1801
· DBLP profile ↗
13ranked-venue papers
11as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TOP-RL: Task-Optimized Progressive Token Pruning with Reinforcement Learning for Vision Language ModelsabstractIn recent years, Large Vision-Language Models (LVLMs) have significantly advanced multimodal tasks. However, their inference requires intensive processing of numerous visual tokens and incurs substantial computational overhead. Existing methods typically compress visual tokens either at the input stage or in early model layers, ignoring variations across tasks and depths. To address these limitations, we introduce TOP-RL, a Task-Optimized Progressive token pruning framework based on Reinforcement Learning. TOP-RL formulates visual token pruning as a multi-stage Markov Decision Process (MDP). It employs an agent trained with dense and fine-grained reward signals to progressively generate differentiable binary masks. This enables TOP-RL to adaptively select crucial visual tokens tailored to each task, effectively balancing accuracy and computational efficiency. Extensive experiments on leading multimodal datasets and advanced LVLMs validate that TOP-RL effectively learns task-optimized pruning policies, significantly boosting inference efficiency while preserving robust performance. For instance, LLaVA-NeXT equipped with TOP-RL achieves a 1.9x speedup in inference time and a 9.3x reduction in FLOPs, with 96% performance preserved. Hengyi Wang, Weiying Xie, Yaotao Wei, Kai Jiang 0001, Mingxiang Cao, Chenhe Hao, Leyuan Fang |
AAAI | 1 |
| 2026 | AuthSig: Safeguarding Scanned Signatures Against Unauthorized Reuse in Paperless WorkflowsabstractWith the deepening trend of paperless workflows, signatures as a means of identity authentication are gradually shifting from traditional ink-on-paper to electronic formats. Despite the availability of dynamic pressure-sensitive and PKI-based digital signatures, static scanned signatures remain prevalent in practice due to their convenience. However, these static images, having almost lost their authentication attributes, cannot be reliably verified and are vulnerable to malicious copying and reuse. To address these issues, we propose AuthSig, a novel static electronic signature framework based on generative models and watermark, which binds authentication information to the signature image. Leveraging the human visual system’s insensitivity to subtle style variations, AuthSig finely modulates style embeddings during generation to implicitly encode watermark bits-enforcing a One Signature, One Use policy. To overcome the scarcity of handwritten signature data and the limitations of traditional augmentation methods, we introduce a keypoint-driven data augmentation strategy that effectively enhances style diversity to support robust watermark embedding. Experimental results show that AuthSig achieves over 98% extraction accuracy under both digital-domain distortions and signature-specific degradations, and remains effective even in print-scan scenarios. Ruiqiang Zhang, Zehua Ma, Guanjie Wang, Chang Liu 0089, Hengyi Wang, Weiming Zhang 0001 |
AAAI | 5 |
| 2025 | 3D Reconstruction with Spatial MemoryabstractWe present Spann3R, a novel approach for dense 3D reconstruction from ordered or unordered image collections. Built on the DUSt3R paradigm, Spann3R uses a transformer-based architecture to directly regress pointmaps from images without any prior knowledge of the scene or camera parameters. Unlike DUSt3R, which pre-dicts per image-pair pointmaps expressed in a local coordinate frame, Spann3R predicts per-image pointmaps expressed in a global coordinate system, thus eliminating the need for optimization-based global alignment. The key idea behind Spann3R is to manage an external spa-tial memory that learns to keep track of all previous relevant 3D information. Spann3R then queries this spatial memory to predict the 3D structure of the next frame in a global coordinate system. Taking advantage of DUSt3R's pre-trained weights, and further fine-tuning on a subset of datasets, Spann3R shows competitive performance and generalization ability on various unseen datasets and can process ordered image collections in real-time. Project page: https://hengyiwang.github.io/projects/spanner Hengyi Wang, Lourdes Agapito |
3DV | 1 |
| 2025 | Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language ModelsabstractHengyi Wang, Haizhou Shi, Shiwei Tan, Weiyi Qin, Wenyuan Wang, Tunyu Zhang, Akshay Nambi, Tanuja Ganu, Hao Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hengyi Wang, Haizhou Shi, Shiwei Tan, Weiyi Qin, Tunyu Zhang, Akshay Uttama Nambi, Tanuja Ganu, Hao Wang 0014 |
NAACL (Long Papers) | 1 |
| 2024 | MorpheuS: Neural Dynamic $360^{\circ}$ Surface Reconstruction from Monocular RGB-D VideoabstractNeural rendering has demonstrated remarkable success in dynamic scene reconstruction. Thanks to the expressiveness of neural representations, prior works can accurately capture the motion and achieve high-fidelity reconstruction of the target object. Despite this, real-world video sce-narios often feature large unobserved regions where neural representations struggle to achieve realistic completion. To tackle this challenge, we introduce MorpheuS, a framework for dynamic$360^\circ$surface reconstruction from a casually captured RGB-D video. Our approach models the target scene as a canonical field that encodes its geometry and appearance, in conjunction with a defor-mation field that warps points from the current frame to the canonical space. We leverage a view-dependent diffusion prior and distill knowledge from it to achieve realistic completion of unobserved regions. Experimental results on various real-world and synthetic datasets show that our method can achieve high-fidelity 360° surface reconstruction of a deformable object from a monocular RGB-D video. Project page: https: / /hengyi wang. gi thub. io/ pro jects/morpheus. Hengyi Wang, Jingwen Wang 0005, Lourdes Agapito |
CVPR | 1 |
| 2024 | Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation ModelsabstractVision transformers (ViTs) have emerged as a significant area of focus, particularly for their capacity to be jointly trained with large language models and to serve as robust vision foundation models. Yet, the development of trustworthy explanation methods for ViTs has lagged, particularly in the context of post-hoc interpretations of ViT predictions. Existing sub-image selection approaches, such as feature-attribution and conceptual models, fall short in this regard. This paper proposes five desiderata for explaining ViTs – faithfulness, stability, sparsity, multi-level structure, and parsimony – and demonstrates the inadequacy of current methods in meeting these criteria comprehensively. We introduce a variational Bayesian explanation framework, dubbed ProbAbilistic Concept Explainers (PACE), which models the distributions of patch embeddings to provide trustworthy post-hoc conceptual explanations. Our qualitative analysis reveals the distributions of patch-level concepts, elucidating the effectiveness of ViTs by modeling the joint distribution of patch embeddings and ViT’s predictions. Moreover, these patch-level explanations bridge the gap between image-level and dataset-level explanations, thus completing the multi-level structure of PACE. Through extensive experiments on both synthetic and real-world datasets, we demonstrate that PACE surpasses state-of-the-art methods in terms of the defined desiderata. Hengyi Wang, Shiwei Tan, Hao Wang 0014 |
ICML | 1 |
| 2024 | FedSLS: Exploring Federated Aggregation in Saliency Latent SpaceabstractFederated Learning (FL) is an emerging direction in distributed machine learning that enables jointly training a global model without sharing data with server. However, data heterogeneity biases the parameter aggregation at the server, leading to slower convergence and poorer accuracy of the global model. To cope with this, most of the existing works involve enforcing regularization in local optimization or improving the model aggregation scheme at the server. Though effective, they lack a deep understanding of cross-client features. In this paper, we propose a saliency latent space feature aggregation method (FedSLS) across federated clients. By Guided BackPropagation (GBP), we transform deep models into powerful and flexible visual fidelity encoders, applicable to general state inputs across different image domains, and achieve powerful aggregation in the form of saliency latent features. Notably, since GBP is label-insensitive, it is sufficient to capture saliency features only once on each client. Experimental results demonstrate that FedSLS leads to significant improvements over the state-of-the-arts in terms of accuracies, especially in highly heterogeneous settings. For example, on CIFAR-10 dataset, FedSLS achieves 63.43% accuracy within the strongly heterogeneous environment α=0.05, which is 6% to 23% higher than other baselines. Hengyi Wang, Weiying Xie, Jitao Ma, Daixun Li, Yunsong Li 0001 |
ACM Multimedia | 1 |
| 2023 | Co-SLAM: Joint Coordinate and Sparse Parametric Encodings for Neural Real-Time SLAMabstractWe present Co-SLAM, a neural RGB-D SLAM system based on a hybrid representation, that performs robust camera tracking and high-fidelity surface reconstruction in real time. Co-SLAM represents the scene as a multi-resolution hash-grid to exploit its high convergence speed and ability to represent high-frequency local features. In addition, Co-SLAM incorporates one-blob encoding, to encourage surface coherence and completion in unobserved areas. This joint parametric-coordinate encoding enables real-time and robust performance by bringing the best of both worlds: fast convergence and surface hole filling. Moreover, our ray sampling strategy allows Co-SLAM to perform global bundle adjustment over all keyframes instead of requiring keyframe selection to maintain a small number of active keyframes as competing neural SLAM approaches do. Experimental results show that Co-SLAM runs at 10-17Hz and achieves state-of-the-art scene reconstruction results, and competitive tracking performance in various datasets and benchmarks (ScanNet, TUM, Replica, Synthetic RGBD). Project page: https://hengyiwang.github.io/projects/CoSLAM Hengyi Wang, Jingwen Wang 0005, Lourdes Agapito |
CVPR | 1 |
| 2022 | Improving Generalization of Deep Networks for Estimating Physical Properties of Containers and FillingsabstractWe present methods to estimate the physical properties of house-hold containers and their fillings manipulated by humans. We use a lightweight, pre-trained convolutional neural network with coordinate attention as a backbone model of the pipelines to accurately locate the object of interest and estimate the physical properties in the CORSMAL Containers Manipulation (CCM) dataset. We address the filling type classification with audio data and then combine this information from audio with video modalities to address the filling level classification. For the container capacity, dimension, and mass estimation, we present a data augmentation and consistency measurement to alleviate the over-fitting issue in the CCM dataset caused by the limited number of containers. We augment the training data using an object-of-interest-based re-scaling that increases the variety of physical values of the containers. We then perform the consistency measurement to choose a model with low prediction variance in the same containers under different scenes, which ensures the generalization ability of the model. Our method improves the generalization ability of the models to estimate the property of the containers that were not previously seen in the training. Hengyi Wang, Chaoran Zhu, Ziyin Ma, Changjae Oh |
ICASSP | 1 |
| 2022 | Boosting Video Object Segmentation Based on Scale InconsistencyabstractWe present a refinement framework to boost the performance of pre-trained semi-supervised video object segmentation (VOS) models. Our work is based on scale inconsistency, which is motivated by the observation that existing VOS models generate inconsistent predictions from input frames with different sizes. We use the scale inconsistency as a clue to devise a pixel-level attention module that aggregates the advantages of the predictions from different-size inputs. The scale inconsistency is also used to regularize the training based on a pixel-level variance measured by an uncertainty estimation. We further present a self-supervised online adaptation, tailored for test-time optimization, that bootstraps the predictions without ground-truth masks based on the scale inconsistency. Experiments on DAVIS 16 and DAVIS 17 datasets show that our framework can be generically applied to various VOS models and improve their performance. Hengyi Wang, Changjae Oh |
ICME | 1 |
| 2021 | Non-Autoregressive Electron Redistribution Modeling for Reaction PredictionabstractReliably predicting the products of chemical reactions presents a fundamental challenge in synthetic chemistry. Existing machine learning approaches typically produce a reaction product by sequentially forming its subparts or intermediate molecules. Such autoregressive methods, however, not only require a pre-defined order for the incremental construction but preclude the use of parallel decoding for efficient computation. To address these issues, we devise a non-autoregressive learning paradigm that predicts reaction in one shot. Leveraging the fact that chemical reactions can be described as a redistribution of electrons in molecules, we formulate a reaction as an arbitrary electron flow and predict it with a novel multi-pointer decoding network. Experiments on the USPTO-MIT dataset show that our approach has established a new state-of-the-art top-1 accuracy and achieves at least 27 times inference speedup over the state-of-the-art methods. Also, our predictions are easier for chemists to interpret owing to predicting the electron flows. Hangrui Bi, Hengyi Wang, Chence Shi, Connor W. Coley, Jian Tang 0005 |
ICML | 2 |
| 2017 | Capacitor voltage regulation of modular multilevel cascaded converter (MMCC-SDBC) as shunt active power filter under different PCC voltagesabstractThis paper presents the application of a modular multilevel cascade converter(MMCC) based on single-delta bridge cells(SDBC) as an active power filter(APF). It is required that the mean value of the DC capacitor voltages is controlled so that the tracking of the source current reference is guaranteed in operation of APF. This paper proposes a DC capacitor voltage regulation method for MMCC-SDBC under different PCC voltages, including the ideal, distorted and unbalanced situations. With the application of the instantaneous symmetrical component theory, this paper analyzes the currents and average active power in three branches of MMCC-SDBC. Based on the analysis and the energy-power relationship of capacitors, the paper develops a systematic regulation method to maintain the DC-bus voltage on a fixed level. The simulation in MATLAB of a three-phase system under different PCC voltage conditions feeding a nonlinear load has shown the effectiveness of the proposed method. Hengyi Wang, Steven Liu |
IECON | 1 |
| 2014 | Integrated current-energy modeling and nonlinear feedback control of modular multilevel STATCOMabstractSTATic synchronous COMpensator (STATCOM) can be integrated into electric transmission systems to provide reactive power compensation and grid voltage support for achieving high-efficient and reliable Flexible AC Transmission Systems (FACTS). An emerging solution for STATCOM with modular structure and multilevel voltage output, named as modular multilevel STATCOM (mmSTATCOM) shows its great advantages of adaption to a wide voltage range, sinusoidal output voltages with less harmonics and less converter losses, compared with the traditional STATCOM. The cascaded Full-Bridge-based mmSTATCOM that is configured with three converter branches into a single delta connection can be classified into the modular multilevel cascade converter (MMCC) family and this configuration can be named as Single-Delta Full-Bridge (SDFB). This paper presents a novel integrated modeling method of SDFB-based STATCOM by taking the independent currents, the total energy and the internal energy balancing into consideration. Then a nonlinear multivariable state-space model is developed. Therefore, a nonlinear control method, named as nonlinear quadratic regulator (NQR), is accordingly proposed to guarantee a safe long-term operation of the SDFB-based STATCOM system both in the balanced and temporarily unbalanced grid conditions. The simulation results are followed to verify the system operations in the normal condition as well as during a temporary grid fault. Hengyi Wang, Jiancheng Tong, Yun Wan, Steven Liu |
IECON | 1 |