EDBT 2026 Demo / reviewers in the wild / expert
Zeyu Hu
dblp:69/10448
· DBLP profile ↗
18ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 9 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MuMA: 3D PBR Texturing via Multi-Channel Multi-View Generation and Albedo Post-ProcessingabstractCurrent methods for 3D generation still fall short in physically based rendering (PBR) texturing, primarily due to limited data and challenges in modeling multi-channel materials. In this work, we propose MuMA, a method for 3D PBR texturing through Multi-channel Multi-view generation and Albedo post-processing. Our approach features two key innovations: 1) we opt to model shaded and albedo appearance channels, where the shaded channels enables the integration intrinsic decomposition modules for material properties; and 2) leveraging multimodal large language models, we emulate artists' techniques for material assessment and selection. Experiments demonstrate that MuMA achieves superior results in visual quality and material fidelity compared to existing methods. Lingting Zhu, Jingrui Ye, Zeyu Hu, Yingda Yin, Lanjiong Li, Jinnan Chen, Shengju Qian, Xin Wang 0178, Qingmin Liao, Lequan Yu |
IEEE Trans. Image Process. | 4 |
| 2025 | MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D GenerationabstractRecent advances in auto-regressive transformers have revolutionized generative modeling across different domains, from language processing to visual generation, demonstrating remarkable capabilities. However, applying these advances to 3D generation presents three key challenges: the unordered nature of 3D data conflicts with sequential next-token prediction paradigm, conventional vector quantization approaches incur substantial compression loss when applied to 3D meshes, and the lack of efficient scaling strategies for higher resolution latent prediction. To address these challenges, we introduce MAR-3D, which integrates a pyramid variational autoencoder with a cascaded masked auto-regressive transformer (Cascaded MAR) for progressive latent upscaling in the continuous space. Our architecture employs random masking during training and auto-regressive denoising in random order during inference, naturally accommodating the unordered property of 3D latent tokens. Additionally, we propose a cascaded training strategy with condition augmentation that enables efficiently up-scale the latent token resolution with fast convergence. Extensive experiments demonstrate that MAR-3D not only achieves superior performance and generalization capabilities compared to existing methods but also exhibits enhanced scaling capabilities compared to joint distribution modeling approaches (e.g., diffusion transformers). Jinnan Chen, Lingting Zhu, Zeyu Hu, Shengju Qian, Yugang Chen, Xin Wang 0178, Gim Hee Lee |
CVPR | 3 |
| 2025 | MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion ParadigmabstractHuman motion generation and editing are key components of computer vision. However, current approaches in this field tend to offer isolated solutions tailored to specific tasks, which can be inefficient and impractical for real-world applications. While some efforts have aimed to unify motion-related tasks, these methods simply use different modalities as conditions to guide motion generation. Consequently, they lack editing capabilities, fine-grained control, and fail to facilitate knowledge sharing across tasks. To address these limitations and provide a versatile, unified framework capable of handling both human motion generation and editing, we introduce a novel paradigm: \textbf{Motion-Condition-Motion}, which enables the unified formulation of diverse tasks with three concepts: source motion, condition, and target motion. Based on this paradigm, we propose a unified framework, \textbf{MotionLab}, which incorporates rectified flows to learn the mapping from source motion to target motion, guided by the specified conditions. In MotionLab, we introduce the 1) MotionFlow Transformer to enhance conditional generation and editing without task-specific modules; 2) Aligned Rotational Position Encoding to guarantee the time synchronization between source motion and target motion; 3) Task Specified Instruction Modulation; and 4) Motion Curriculum Learning for effective multi-task learning and knowledge sharing across tasks. Notably, our MotionLab demonstrates promising generalization capabilities and inference efficiency across multiple benchmarks for human motion. Our code and additional video results are available at: https://diouo.github.io/motionlab.github.io/. Ziyan Guo, Zeyu Hu, De Wen Soh, Na Zhao 0004 |
ICCV | 2 |
| 2025 | Contrastive Learning Method for Behavior Prediction and Sequential Recommendation based on Multi-Intention DisentanglementabstractSequential recommendation is one of the important branches of recommender system, aiming to achieve personalized recommended items for the future through the analysis and prediction of users’ ordered historical interactive behaviors. However, along with the growth of the user volume and the increasingly rich behavioral information, how to understand and disentangle the user’s interactive multi-intention effectively also poses challenges to behavior prediction and sequential recommendation. In light of these challenges, we propose a Contrastive Learning sequential recommendation method based on Multi-Intention Disentanglement (MIDCL). In our work, intentions are recognized as dynamic and diverse, and user behaviors are often driven by current multi-intentions, which means that the model needs to not only mine the most relevant implicit intention for each user, but also impair the influence from irrelevant intentions. Therefore, we choose Variational Auto-Encoder (VAE) to realize the disentanglement of users’ multi-intentions. We propose two types of contrastive learning paradigms for finding the most relevant user’s interactive intention, and maximizing the mutual information of positive sample pairs, respectively. Experimental results show that MIDCL not only has significant superiority over most existing baseline methods, but also brings a more interpretable case to the research about intention-based prediction and recommendation. Zeyu Hu, Yuzhi Xiao, Xuanrong Huo |
IJCNN | 1 |
| 2025 | Light-SQ: Structure-aware Shape Abstraction with Superquadrics for Generated MeshesabstractIn user-generated-content (UGC) applications, non-expert users often rely on image-to-3D generative models to create 3D assets. In this context, primitive-based shape abstraction offers a promising solution for UGC scenarios by compressing high-resolution meshes into compact, editable representations. Towards this end, effective shape abstraction must therefore be structure-aware, characterized by low overlap between primitives, part-aware alignment, and primitive compactness. We present Light-SQ, a novel superquadric-based optimization framework that explicitly emphasizes structure-awareness from three aspects. (a) We introduce SDF carving to iteratively udpate the target signed distance field, discouraging overlap between primitives. (b) We propose a block-regrow-fill strategy guided by structure-aware volumetric decomposition, enabling structural partitioning to drive primitive placement. (c) We implement adaptive residual pruning based on SDF update history to surpress over-segmentation and ensure compact results. In addition, Light-SQ supports multiscale fitting, enabling localized refinement to preserve fine geometric details. To evaluate our method, we introduce 3DGen-Prim, a benchmark extending 3DGen-Bench with new metrics for both reconstruction quality and primitive-level editability. Extensive experiments demonstrate that Light-SQ enables efficient, high-fidelity, and editable shape abstraction with superquadrics for complex generated geometry, advancing the feasibility of 3D UGC creation. Project Page: https://johann.wang/Light-SQ/ . Yuhan Wang 0002, Weikai Chen 0001, Zeyu Hu, Yingda Yin, Keyang Luo, Shengju Qian, Yiyan Ma, Yuhuan Zhou, Hao Luo 0001, Wan Wang, Xiaobin Shen 0004, Kuixin Zhu, Chuanlang Hong, Lijie Feng, Xin Wang 0178, Chen Change Loy |
SIGGRAPH Asia | 3 |
| 2025 | Analysis of Preamble Collisions in Satellite-Based Internet of Things Using Static Floor FieldabstractRandom access protocols play a pivotal role in the development of Satellite-based Internet of Things (S-IoT). However, existing research often inadequately addresses the rapid increase in device numbers, leading to theoretical models that unconvincingly simulate group behavior based on single-device behavior. Instead, we employ the static floor field method, traditionally used to analyze the evacuation of large numbers of pedestrians from confined spaces, to investigate the competition dynamics for limited preambles among numerous devices within the S-IoT system. This approach offers significant advantages in capturing dynamics and ensuring scalability during massive interactive sessions. Specifically, we extend the analysis model, based on the static floor field method, to address the issue of preamble collisions among a large number of IoT devices. The data obtained from this method closely align with the results from a system-level simulation platform, substantiating the accuracy and effectiveness of our analytical approach. Consequently, the analytical and predictive capabilities of the proposed method lead to the development of intelligent preamble resource allocation strategies based on reinforcement learning, thereby enhancing the overall performance of the S-IoT system. The efficacy of the proposed strategy is further validated by results derived from a simulation platform using authentic orbital data. Zeyu Hu, Zhiyuan Jiang |
IEEE Internet Things J. | 1 |
| 2025 | Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation of Indoor ScenesabstractIn recent years, sparse voxel-based methods have become the state-of-the-arts for 3D semantic segmentation of indoor scenes, thanks to the powerful 3D CNNs. Nevertheless, being oblivious to the underlying geometry, voxel-based methods suffer from ambiguous features on spatially close objects and struggle with handling complex and irregular geometries due to the lack of geodesic information. In view of this, we present Voxel-Mesh Network (VMNet), a novel 3D deep architecture that operates on the voxel and mesh representations leveraging both the euclidean and geodesic information. Intuitively, the euclidean information extracted from voxels can offer contextual cues representing interactions between nearby objects, while the geodesic information extracted from meshes can help separate objects that are spatially close but have disconnected surfaces. To incorporate such information from the two domains, we design an intra-domain attentive module for effective feature aggregation and an inter-domain attentive module for adaptive feature fusion. Experimental results validate the effectiveness of VMNet: specifically, on the challenging ScanNet dataset for large-scale segmentation of indoor scenes, it outperforms the state-of-the-art SparseConvNet and MinkowskiNet (74.6% versus 72.5% and 73.6% in mIoU) with a simpler network structure (17M versus 30M and 38M parameters). Zeyu Hu, Xuyang Bai, Jiaxiang Shang, Jiayu Dong, Xin Wang 0178, Guangyuan Sun, Hongbo Fu 0001, Chiew-Lan Tai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Channel Knowledge Map-Aided Channel Prediction With Measurements-Based EvaluationabstractGaining accurate channel state information (CSI) through a low-cost scheme has always been difficult in wireless communication systems. One of the current research directions is to obtain the CSI from the channel knowledge map (CKM) based on the users’ location. However, the direct utilization of CSI in CKM is hindered due to the sensitivity of instantaneous CSI to time-varying scattering environments and positioning errors. To address this issue, this paper proposes a channel prediction scheme that combines the CKM with historical user CSI to enhance the beamforming performance in multiple-input multiple-output (MIMO) systems. Specifically, the joint-orthogonal matching pursuit algorithm is used to accurately reconstruct the user channel with high precision using a limited number of pilots, and the multi-path components tracking algorithm is employed to extract the common and independent support sets of paths from the estimated channel and the CKM. Lastly, an adaptive and low-complexity predictor is utilized to obtain the future user CSI. The proposed scheme has been evaluated using multiple measured channel datasets, the results indicate a significant improvement in predicting channel cosine similarity compared to directly using the CSI from CKM and existing schemes. Xianling Wang, Yi Shi 0004, Yingyujiao Huang, Zeyu Hu, Lin Chen 0037, Zhiyuan Jiang |
IEEE Trans. Commun. | 5 |
| 2024 | Can Channel Knowledge Map Help to Predict Instantaneous MIMO Channel State Information?abstractAccurate channel state information (CSI) is crucial for optimizing system performance in wireless communications. One of the current research directions is to extract the CSI from the channel knowledge map (CKM) based on users' location. While such an approach is promising for large-scale CSI acquisition, e.g., pathloss, it is still questionable whether CKM can help to predict instantaneous CSI which is very sensitive to positioning error and scattering environment changes. To answer this question, this paper proposes a channel prediction scheme that combines the CKM with historical user CSI to improve the beamforming performance in multiple-input multiple-output (MIMO) systems. Specifically, it utilizes the complex amplitude information of multi-path components (MPCs) and employs a low-complexity method to predict the future MIMO CSI. Through experiments based on real-world channel data, the results demonstrate that the proposed scheme outperforms state-of-the-art ones that use the CKM alone or autoregressive schemes without CKM. In the non-line-of-sight scenarios, when the positioning error exceeds 0.525 m, small-scale CSI in CKM provides little gain. In line-of-sight environments, the threshold for the usability of small-scale CSI in terms of positioning error is approximately 0.725 m or slightly higher. When the positioning error is less than 0.225 m, the CKM is beneficial to predict instantaneous CSI in both scenarios. Xianling Wang, Yi Shi 0004, Yingyujiao Huang, Zeyu Hu, Zhiyuan Jiang |
WCNC | 5 |
| 2024 | A Fluid Limit Approach to Age of Information Optimization in Multiaccess Networks With Transmission Frequency ConstraintsabstractThis paper examines the use of the fluid limit, a mathematical tool used to analyze Age of Information (AoI) in networks, to study the long-term average AoI for scheduling optimization in multiaccess networks with transmission frequency constraints. Two types of problems that have been studied in the literature but remain unsolved are revisited under this approach, wherein multiple agents transmit over an error-prone multiaccess channel, and long-term transmission frequency constraints are considered with either a resource constraint or a minimum throughput requirement. Previous works have derived the Whittle’s Index (WI) policy, max-weight policy and average AoI lower bound, however without optimality guarantee or closed-form performance analysis. This work advances the field by deriving closed-form optimal AoI and achieving scheduling policies for both problems, by utilizing the fluid limit tool and transforming the original high-dimensional Markov decision process into solving a set of partial derivative equations. As a result, threshold-based scheduling policies with closed-form threshold expressions are obtained which can be proven to be asymptotically optimal when the number of agents is large. Numerical simulations are conducted to demonstrate the performance optimality of the proposed schemes and analytical results. Zeyu Hu, Zhiyuan Jiang |
IEEE Trans. Commun. | 2 |
| 2022 | TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersabstractLiDAR and camera are two important sensors for 3D object detection in autonomous driving. Despite the increasing popularity of sensor fusion in this field, the robustness against inferior image conditions, e.g., bad illumination and sensor misalignment, is under-explored. Existing fusion methods are easily affected by such conditions, mainly due to a hard association of LiDAR points and image pixels, established by calibration matrices. We propose TransFusion, a robust solution to LiDAR-camera fusion with a soft-association mechanism to handle inferior image conditions. Specifically, our TransFusion consists of convolutional backbones and a detection head based on a transformer decoder. The first layer of the decoder predicts initial bounding boxes from a LiDAR point cloud using a sparse set of object queries, and its second decoder layer adaptively fuses the object queries with useful image features, leveraging both spatial and contextual relationships. The attention mechanism of the transformer enables our model to adaptively determine where and what information should be taken from the image, leading to a robust and effective fusion strategy. We additionally design an image-guided query initialization strategy to deal with objects that are difficult to detect in point clouds. TransFusion achieves state-of-the-art performance on large-scale datasets. We provide extensive experiments to demonstrate its robustness against degenerated image quality and calibration errors. We also extend the proposed method to the 3D tracking task and achieve the 1st place in the leader-board of nuScenes tracking, showing its effectiveness and generalization capability. [code release] Xuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang, Hongbo Fu 0001, Chiew-Lan Tai |
CVPR | 2 |
| 2022 | LiDAL: Inter-frame Uncertainty Based Active Learning for 3D LiDAR Semantic Segmentation
Zeyu Hu, Xuyang Bai, Xin Wang 0178, Guangyuan Sun, Hongbo Fu 0001, Chiew-Lan Tai |
ECCV (27) | 1 |
| 2022 | Energy Saving in LEO-HTS Constellation Based on Adaptive Power Allocation with Multi-Beam Directivity ControlabstractHigh throughput satellite (HTS) has been identified as a key technology for improving communication capacity of satellite system. Recently, low earth orbit high throughput satellite (LEO-HTS) has attracted much attention because of its low latency and construction cost. However, LEO-HTS is power-limited when it passes through the shadow of the earth, which will affect the lifetime and capabilities of the satellite. This paper focuses on the problem of energy saving for multiple LEO-HTSs in the LEO-HTS constellation. A method of adaptive power allocation with multi-beam directivity control according to the traffic demand of user equipments (UEs) is proposed to save energy. In this case, an optimization model is established to minimize the transmission power of multiple LEO-HTSs. Then, an alternating optimization framework is introduced to iteratively solve the transmission power and multi-beam directivity of multiple LEO-HTSs. Finally, numerical results demonstrate that the proposed method can save energy more effectively than the traditional method in the scenario with high traffic demand and concentrated UE distribution. Jinming Zhao, Yong Li 0001, Zeyu Hu |
PIMRC | 3 |
| 2021 | PointDSC: Robust Point Cloud Registration Using Deep Spatial ConsistencyabstractRemoving outlier correspondences is one of the critical steps for successful feature-based point cloud registration. Despite the increasing popularity of introducing deep learning techniques in this field, spatial consistency, which is essentially established by a Euclidean transformation between point clouds, has received almost no individual attention in existing learning frameworks. In this paper, we present PointDSC, a novel deep neural network that explicitly incorporates spatial consistency for pruning outlier correspondences. First, we propose a nonlocal feature aggregation module, weighted by both feature and spatial coherence, for feature embedding of the input correspondences. Second, we formulate a differentiable spectral matching module, supervised by pairwise spatial compatibility, to estimate the inlier confidence of each correspondence from the embedded features. With modest computation cost, our method outperforms the state-of-the-art hand- crafted and learning-based outlier rejection approaches on several real-world datasets by a significant margin. We also show its wide applicability by combining PointDSC with different 3D local descriptors. [code release] Xuyang Bai, Zixin Luo, Lei Zhou 0011, Lei Li 0038, Zeyu Hu, Hongbo Fu 0001, Chiew-Lan Tai |
CVPR | 6 |
| 2021 | Learning to Match Features with Seeded Graph Matching NetworkabstractMatching local features across images is a fundamental problem in computer vision. Targeting towards high accuracy and efficiency, we propose Seeded Graph Matching Network, a graph neural network with sparse structure to reduce redundant connectivity and learn compact representation. The network consists of 1) Seeding Module, which initializes the matching by generating a small set of reliable matches as seeds. 2) Seeded Graph Neural Network, which utilizes seed matches to pass messages within/across images and predicts assignment costs. Three novel operations are proposed as basic elements for message passing: 1) Attentional Pooling, which aggregates keypoint features within the image to seed matches. 2) Seed Filtering, which enhances seed features and exchanges messages across images. 3) Attentional Unpooling, which propagates seed features back to original keypoints. Experiments show that our method reduces computational and memory complexity significantly compared with typical attention-based networks while competitive or higher performance is achieved. Zixin Luo, Lei Zhou 0011, Xuyang Bai, Zeyu Hu, Chiew-Lan Tai, Long Quan |
ICCV | 6 |
| 2021 | VMNet: Voxel-Mesh Network for Geodesic-Aware 3D Semantic SegmentationabstractIn recent years, sparse voxel-based methods have be-come the state-of-the-arts for 3D semantic segmentation of indoor scenes, thanks to the powerful 3D CNNs. Nevertheless, being oblivious to the underlying geometry, voxel-based methods suffer from ambiguous features on spatially close objects and struggle with handling complex and irregular geometries due to the lack of geodesic information. In view of this, we present Voxel-Mesh Network (VMNet), a novel 3D deep architecture that operates on the voxel and mesh representations leveraging both the Euclidean and geodesic information. Intuitively, the Euclidean information extracted from voxels can offer contextual cues representing interactions between nearby objects, while the geodesic information extracted from meshes can help separate objects that are spatially close but have disconnected surfaces. To incorporate such information from the two domains, we design an intra-domain attentive module for effective feature aggregation and an inter-domain attentive module for adaptive feature fusion. Experimental results validate the effectiveness of VMNet: specifically, on the challenging ScanNet dataset for large-scale segmentation of indoor scenes, it outperforms the state-of-the-art SparseConvNet and MinkowskiNet (74.6% vs 72.5% and 73.6% in mIoU) with a simpler network structure (17M vs 30M and 38M parameters). Code release: https://github.com/hzykent/VMNet Zeyu Hu, Xuyang Bai, Jiaxiang Shang, Jiayu Dong, Xin Wang 0178, Guangyuan Sun, Hongbo Fu 0001, Chiew-Lan Tai |
ICCV | 1 |
| 2020 | JSENet: Joint Semantic Segmentation and Edge Detection Network for 3D Point Clouds
Zeyu Hu, Mingmin Zhen, Xuyang Bai, Hongbo Fu 0001, Chiew-Lan Tai |
ECCV (20) | 1 |
| 2020 | Uncertain Gompertz regression model with imprecise observations
Zeyu Hu |
Soft Comput. | 1 |