EDBT 2026 Demo / reviewers in the wild / expert
Feng Wu 0005
dblp:25/3972-5
· DBLP profile ↗
23ranked-venue papers
0as first author
23since 2021 · last 2026
0000-0001-7266-5579ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniSOT: A Unified Framework for Multi-Modality Single Object TrackingabstractSingle object tracking aims to localize target object with specific reference modalities (bounding box, natural language or both) in a sequence of specific video modalities (RGB, RGB+Depth, RGB+Thermal or RGB+Event.). Different reference modalities enable various human-machine interactions, and different video modalities are demanded in complex scenarios to enhance tracking robustness. Existing trackers are designed for single or several video modalities with single or several reference modalities, which leads to separate model designs and limits practical applications. Practically, a unified tracker is needed to handle various requirements. To the best of our knowledge, there is still no tracker that can perform tracking with these above reference modalities across these video modalities simultaneously. Thus, in this paper, we present a unified tracker, UniSOT, for different combinations of three reference modalities and four video modalities with uniform parameters. Extensive experimental results on 18 visual tracking, vision-language tracking and RGB+X tracking benchmarks demonstrate that UniSOT shows superior performance against modality-specific counterparts. Notably, UniSOT outperforms previous counterparts by over 3.0% AUC on TNL2K across all three reference modalities and outperforms Un-Track by over 2.0% main metric across all three RGB+X video modalities. Yinchao Ma, Yuyang Tang 0001, Wenfei Yang, Tianzhu Zhang 0001, Feng Wu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | NVC-1B: Scaling up Neural Video Coding ModelsabstractEmerging large models have achieved notable progress in the fields of natural language processing and computer vision. However, large models for neural video coding are still unexplored. In this paper, we try to explore how to build a large neural video coding model. Based on a small baseline model, we gradually scale up the model sizes of its different coding parts, including the motion encoder-decoder, motion entropy model, contextual encoder-decoder, contextual entropy model, and temporal context mining module, and analyze the influence of model sizes on video compression performance. Then, we explore using different architectures, including CNN, mixed CNN-Transformer, and Transformer architectures, to implement the neural video coding model and analyze the influence of model architectures on video compression performance. Based on our exploration results, we design the first neural video coding model having more than 1 billion parameters - NVC-1B. Experimental results show that our large model achieves a significant video compression performance improvement over recent state-of-the-art neural video compression models. With the continuous advancement in hardware and the successful on-device deployment of large models, we anticipate that our proposed large neural video coding model can bring video coding technologies to the next level. Chuanbo Tang, Xihua Sheng, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | DA2-LiDAR: A Generic Density-Adaptive Framework for Unsupervised Domain Adaptation in LiDAR SegmentationabstractThis paper addresses the critical challenge of domain adaptation for LiDAR-based semantic segmentation, particularly the significant density disparities that emerge when transferring models from synthetic to real-world environments. We present DA2-LiDAR, a novel density-adaptive domain adaptation framework that bridges domain gaps through the construction of intermediate domains with density-varying point distributions. Our approach employs a simple yet effective masking strategy that systematically reduces density discrepancies between domains while extracting more effective supervisory signals, as well as preserving critical semantic information. The framework consists of three key components: (1) a Density Adaptation Module that establishes a continuous spectrum of intermediate domains through dataset-agnostic masking operations; (2) a Contextual Consistency Module that enforces relational coherence across differently masked variants of the same scan at varying degrees, providing additional supervision signals, enhancing the model's ability to extract features; and (3) a Semantic Preservation Module that mitigates information loss in heavily masked scans by reconstructing domain-specific data distributions. Extensive experiments on synthetic-to-real and other benchmarks demonstrate that DA2-LiDAR consistently outperforms state-of-the-art methods, achieving significant improvements in cross-domain generalization without requiring dataset-specific prior knowledge or introducing computational overhead. Rui Sun 0006, Wangkai Li, Naisong Luo, Yuan Wang 0064, Tianzhu Zhang 0001, Feng Wu 0005 |
IEEE Trans. Image Process. | 7 |
| 2026 | TVRN: Invertible Neural Networks for Compression-Aware Temporal Video RescalingabstractTo fit diverse display and bandwidth constraints, high-frame-rate videos are temporally downscaled to low-frame-rate (LFR) and later upscaled, requiring joint optimization for effective frame-rate rescaling. However, existing methods typically link the two operations via training objectives, without fully exploiting their reciprocal nature, which may cause high-frequency information loss. Moreover, they overlook the impact of lossy codecs on LFR videos, limiting real-world applicability. In this work, we propose an end-to-end framework for compression-aware frame-rate rescaling, named TVRN. To regularize high-frequency information lost during frame-rate downscaling, TVRN adopts an invertible architecture that combines a Multi-Input Multi-Output Temporal Wavelet Transform with a high-frequency reconstruction module. To enable end-to-end training through non-differentiable lossy codecs, we design a surrogate network that approximates their gradients. Finally, to improve robustness under various compression levels, we extend TVRN to an asymmetric architecture by incorporating compression-aware features learned via a learning-to-rank strategy. Extensive experiments show that TVRN outperforms existing methods in reconstruction quality under industrial video compression settings. Source code is publicly available at https://github.com/fengxinmin/TVRN_public. Xinmin Feng, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Image Process. | 4 |
| 2026 | Neuron Segment Connectivity Prediction With Multimodal Features for ConnectomicsabstractReconstructing neurons from large electron microscopy (EM) datasets for connectomic analysis presents a significant challenge, particularly in segmenting neurons of complex morphologies. Previous deep learning-based neuron segmentation methods often rely on pixel-level image context and produce extensive oversegmented fragments. Detecting these split errors and merging the split neuron segments are non-trivial for various neurons in a large-scale EM data volume. In this work, we exploit multimodal features in the full workflow of automatic neuron proofreading. We propose a novel connection point detection network that utilizes both global 3D morphological features and high-resolution local image context to extract candidate segment pairs from massive adjacent segments. To effectively fuse the 3D morphological feature and the dense image features from very different scales, we design a proposal-based image feature sampling to improve the efficiency of multimodal cross-attentions. Integrating the connection point detection network with our connectivity prediction network which also utilizes multimodal features, we make a fully automatic neuron segment merging pipeline, closely imitating human proofreading. Comprehensive experimental results verify the effectiveness of the proposed modules and demonstrate the robustness of the entire pipeline in large-scale neuron reconstruction. The code and data are available at https://github.com/Levishery/Neuron-Segment-Connection-Prediction. Qihua Chen, Xuejin Chen, Chenxuan Wang, Zhiwei Xiong, Feng Wu 0005 |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Partition Map-Based Fast Block Partitioning for VVC Inter CodingabstractAmong the new techniques of Versatile Video Coding (VVC), the quadtree with nested multi-type tree (MTT) block structure yields significant coding gains by providing more flexible block partitioning patterns. However, the recursive partition search in the VVC encoder increases the encoder complexity substantially. To address this issue, we propose a partition map-based algorithm to pursue fast block partitioning in inter coding. Based on our previous work on partition map-based methods for intra coding, we analyze the characteristics of VVC inter coding and improve the partition map by incorporating an MTT mask for early termination. Next, we develop a neural network that uses both spatial and temporal features to predict the partition map. It consists of several special designs, including stacked top-down and bottom-up processing, quantization parameter modulation layers, and partitioning-adaptive warping. Furthermore, we present a dual-threshold decision scheme to achieve a fine-grained trade-off between complexity reduction and rate-distortion performance loss. The experimental results demonstrate that the proposed method achieves an average 51.30% encoding time saving with a 2.12% Bjøntegaard-delta-bit-rate under the random access configuration. The source code is publicly available athttps://github.com/ustc-ivclab/IPM. Xinmin Feng, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Multim. | 5 |
| 2026 | USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020sabstractImage/video coding has been a remarkable research area for both academia and industry for many years. Testing datasets, especially high-quality image/video datasets, are desirable for the justified evaluation of coding-related research, practical applications, and standardization activities. We put forward a test dataset, namely USTC-TD, which has been successfully adopted in the practical end-to-end image/video coding challenge ofIEEE International Conference on Visual Communications and Image Processing (VCIP)in 2022 and 2023. USTC-TD contains 40 images at 4K spatial resolution and 10 video sequences at 1080p spatial resolution, featuring various content due to the diverse environmental factors (e.g., scene type, texture, motion, view) and the designed imaging factors (e.g., illumination, lens, shadow). We quantitatively evaluate USTC-TD on different image/video features (spatial, temporal, color, lightness), and compare it with the previous image/video test datasets, which verifies its excellent compensation for the shortcomings of existing datasets. We also evaluate both classic standardized and recently learned image/video coding schemes on USTC-TD using objective quality metrics (PSNR, MS-SSIM, VMAF) and subjective quality metric (MOS), providing an extensive benchmark for these evaluated schemes. Based on the characteristics and specific design of the proposed test dataset, we analyze the benchmark performance and shed light on the future research and development of image/video coding. All the data are released online:https://esakak.github.io/USTC-TD. Zhuoyuan Li 0001, Junqi Liao, Chuanbo Tang, Haotian Zhang 0009, Yifan Bian, Xihua Sheng, Xinmin Feng, Yao Li 0016, Changsheng Gao, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Multim. | 13 |
| 2025 | Structural and Statistical Texture Knowledge Distillation and Learning for SegmentationabstractLow-level texture feature/knowledge is also of vital importance for characterizing the local structural pattern and global statistical properties, such as boundary, smoothness, regularity, and color contrast, which may not be well addressed by high-level deep features. In this paper, we aim to re-emphasize the low-level texture information in deep networks for semantic segmentation and related knowledge distillation tasks. To this end, we take full advantage of both structural and statistical texture knowledge and propose a novel Structural and Statistical Texture Knowledge Distillation (SSTKD) framework for semantic segmentation. Specifically, Contourlet Decomposition Module (CDM) is introduced to decompose the low-level features with iterative Laplacian pyramid and directional filter bank to mine the structural texture knowledge, and Texture Intensity Equalization Module (TIEM) is designed to extract and enhance the statistical texture knowledge with the corresponding Quantization Congruence Loss (QDL). Moreover, we propose the Co-occurrence TIEM (C-TIEM) and generic segmentation frameworks, namely STLNet++ and U-SSNet, to enable existing segmentation networks to harvest the structural and statistical texture information more effectively. Extensive experimental results on three segmentation tasks demonstrate the effectiveness of the proposed methods and their state-of-the-art performance on seven popular benchmark datasets, respectively. Deyi Ji, Feng Zhao 0004, Hongtao Lu 0001, Feng Wu 0005, Jieping Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Long-Term Feature Extraction via Frequency Prediction for Efficient Reinforcement LearningabstractSample efficiency remains a key challenge for the deployment of deep reinforcement learning (RL) in real-world scenarios. A common approach is to learn efficient representations through future prediction tasks, facilitating the agent to make farsighted decisions that benefit its long-term performance. Existing methods extract predictive features by predicting multi-step future state signals. However, they do not fully exploit the structural information inherent in sequential state signals, which can potentially improve the quality of long-term decision-making but is difficult to discern in the time domain. To tackle this problem, we introduce a new perspective that leverages the frequency domain of state sequences to extract the underlying patterns in time series data. We theoretically show that state sequences contain structural information closely tied to policy performance and signal regularity and analyze the fitness of the frequency domain for extracting these two types of structural information. Inspired by that, we propose a novel representation learning method, State Sequences Prediction via Fourier Transform (SPF), which extracts long-term features by predicting the Fourier transform of infinite-step future state sequences. The appealing features of our frequency prediction objective include: 1) simple to implement due to a recursive relationship; 2) providing an upper bound on the performance difference between the optimal policy and the latent policy in the representation space. Experiments on standard and goal-conditioned RL tasks demonstrate that the proposed method outperforms several state-of-the-art algorithms in terms of both sample efficiency and performance. Jie Wang 0005, Mingxuan Ye, Yufei Kuang, Rui Yang 0031, Wengang Zhou 0001, Houqiang Li, Feng Wu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Potential Field-Based and Network State-Aware Anycast Routing for LEO Satellite NetworksabstractLow Earth Orbit Satellite Networks (LEOSNs) have emerged as a promising paradigm for space information networks, where multiple inter-satellite links facilitate the rapid transmission of on-orbit data to multi-ground station systems. When the destination of the transmission is not a certain ground station, but anyone of the ground stations, it can be modeled as an anycast problem. However, the time-varying topology, dynamic inter-satellite link status and limited onboard resources bring challenges to the routing of on-orbit data. Existing unicast routing solutions failed to address the anycast routing problem as they could not fully utilize multiple ground stations. Inspired by the Potential Field (PF) theory in physics, we make the first attempt to adopt the PF approach in the on-orbit data anycast routing problem. We design a synthesized PF model including several sub-fields corresponding to network states such as length of path, node transmission load and link bandwidth. By implementing inter-satellite propagation and synthesis of PF, dynamic perception and unified measurement of network states can be achieved. Based on our PF model, we propose a distributed PF-based and Network State-aware Anycast Routing (PFNSAR) algorithm, which regards the ground stations as multiple sources of attractive potential and guides packets along the gradient of synthesized PF. Meanwhile, the occurrence of the well-known local minimum is novelly handled by setting PF configuration subject to a parameter constraint and switching routing modes when routing packets. Extensive simulations demonstrate that PFNSAR provides unified assessment of multi-dimensional network states, reduces up to 30% of average delivery delay, and improves the performance including packet delivery rate and load balancing compared with existing works. Guangyuan Wei, Yunpeng Hou, Shuangwu Chen, Jian Yang 0014, Huasen He, Feng Wu 0005 |
IEEE Trans. Commun. | 6 |
| 2025 | Semantic-Aware Late-Stage Supervised Contrastive Learning for Fine-Grained Action RecognitionabstractFine-grained action recognition typically faces challenges with lower inter-class variances and higher intra-class variances. Supervised contrastive learning is inherently suitable for this task, as it can decrease intra-class feature distances while increasing inter-class ones. However, directly applying it into fine-grained action recognition encounters two main problems. The first problem stems from the heavy training cost associated with supervised contrastive learning, which requires numerous training epochs, each involving double augmentation views per instance. To address this issue, we propose the late-stage supervised contrastive learning (late-SC) strategy, which effectively reduces the number of training epochs needed for the contrastive learning process. The second problem is that supervised contrastive loss does not explicitly consider the semantic distances between fine-grained actions when adjusting representation distances. This results in less reasonable and efficient adjustments to the representation space. To overcome this limitation, we introduce the semantic-aware temperature adaptation (STA) mechanism, enhancing the suitability of the supervised contrastive loss for fine-grained action recognition. We conduct experiments on several benchmark datasets for fine-grained action recognition, including Epic-Kitchens-55/100, SomethingSomething-V1, and Diving48-V2. The results demonstrate that our proposed method (referred to as LSC-STA) consistently enhances performance across various base feature extractors, without introducing additional inference overhead and incurring only a marginal increase in training expenses. Yijun Pan, Yueyi Zhang 0001, Zilei Wang, Xiaoyan Sun 0001, Feng Wu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Deep Ensemble Stochastic Configuration Network via Graph Intuitionistic Fuzzy for Depression RecognitionabstractDepression is an affective disorder that poses a serious threat to both mental and physical health. Utilizing fuzzy -based neural network models for the identification and screening of depression can facilitate early intervention and treatment. Intuitionistic fuzzy stochastic configuration networks (IFSCNs) utilize a cost-sensitive learning framework, which enhances generalization performance for solving binary classification problems. However, IFSCNs ignore the impact of the relative neighborhood density of imbalanced samples with outliers. To learn more discriminant information from class imbalance depression recognition task, in this paper, we propose a novel deep ensemble stochastic configuration network via graph intuitionistic fuzzy, termed as DeSCN-GIF. Specifically, we first use graph-based intuitionistic fuzzy method to determine the membership and non-membership functions through the relative neighborhood density of imbalanced samples; moreover, an incremental self-ensemble deep stochastic configuration framework is presented to learn multi-level discriminative features, in which graph intuitionistic fuzzy cost-sensitive least squares loss function and weighted supervision mechanisms are applied to determine the parameters of DeSCN-GIF. Experimental results on Chinese syllable-based imbalanced depression voice datasets show that DeSCN-GIF has better binary classification performance compared to other learning models such as IFSCN, DSCN, SCN, GE-IFRVFL-CIL, IFRVFL, EDRVFL, DRVFL, RVFL, and DIFL-TSVM. Chenglong Zhang 0001, Dawei Cheng, Jiankai Xue, Feng Wu 0005, David Zhang 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2025 | Robust Decorrelated Stochastic Configuration Networks Ensemble via Weighted Negative Correlation LearningabstractStochastic configuration network (SCN) is a kind of incremental random neural network that assigns input weights and biases through data-dependent supervisory mechanism. However, the robustness of SCN is significantly reduced when processing the data disturbed by outliers. Aiming at improve the noisy data regression performance of SCN, this article presents a novel robust decorrelated SCNs ensemble model (RDSCNE). Such a robust decorrelated ensemble framework adopts weighted negative correlation learning (WNCL) and a robust regularization technique, which can guarantee the generalization performance for noisy data processing. Specifically, we first present a WNCL framework based on kernel density estimation (KDE) to build SCNs ensemble model, so that the negative effects of noise can be suppressed through KDE to calculate penalty weights of each training sample for the computation of ensemble weights. Meanwhile,l1norm loss function combined withl2regularization technique is employed as the objective function of base components. This approach is designed to process outliers with sparse characteristics and alleviate the over-fitting phenomenon. Then, augmented Lagrange multiplier (ALM) method is used to calculate the objective function. Experimental results over some regression datasets with Gaussian outliers demonstrate that the proposed RDSCNE model has better robustness than the various SCN variants. Chenglong Zhang 0001, Chaoxun Guo, Shifei Ding, Feng Wu 0005, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2024 | Scene Adaptive Sparse Transformer for Event-based Object DetectionabstractWhile recent Transformer-based approaches have shown impressive performances on event-based object detection tasks, their high computational costs still diminish the low power consumption advantage of event cameras. Image-based works attempt to reduce these costs by introducing sparse Transformers. However, they display inade-quate sparsity and adaptability when applied to event-based object detection, since these approaches cannot balance the fine granularity of token-level sparsification and the efficiency of window-based Transformers, leading to re-duced performance and efficiency. Furthermore, they lack scene-specific sparsity optimization, resulting in information loss and a lower recall rate. To overcome these limi-tations, we propose the Scene Adaptive Sparse Transformer (SAST). SAST enables window-token co-sparsification, sig-nificantly enhancing fault tolerance and reducing compu-tational overhead. Leveraging the innovative scoring and selection modules, along with the Masked Sparse Window Self-Attention, SAST showcases remarkable scene-aware adaptability: It focuses only on important objects and dy-namically optimizes sparsity level according to scene complexity, maintaining a remarkable balance between performance and computational cost. The evaluation results show that SAST outperforms all other dense and sparse networks in both performance and efficiency on two large-scale event-based object detection datasets (1 Mpx and Genl). Code: https://github.com/Peterande/SAST. Yansong Peng, Hebei Li, Yueyi Zhang 0001, Xiaoyan Sun 0001, Feng Wu 0005 |
CVPR | 5 |
| 2024 | Label Deconvolution for Node Representation Learning on Large-Scale Attributed Graphs Against Learning BiasabstractNode representation learning on attributed graphs-whose nodes are associated with rich attributes (e.g., texts and protein sequences)-plays a crucial role in many important downstream tasks. To encode the attributes and graph structures simultaneously, recent studies integrate pre-trained models with graph neural networks (GNNs), where pre-trained models serve as node encoders (NEs) to encode the attributes. As jointly training large NEs and GNNs on large-scale graphs suffers from severe scalability issues, many methods propose to train NEs and GNNs separately. Consequently, they do not take feature convolutions in GNNs into consideration in the training phase of NEs, leading to a significant learning bias relative to the joint training. To address this challenge, we propose an efficient label regularization technique, namely Label Deconvolution (LD), to alleviate the learning bias by a novel and highly scalable approximation to the inverse mapping of GNNs. The inverse mapping leads to an objective function that is equivalent to that by the joint training, while it can effectively incorporate GNNs in the training phase of NEs against the learning bias. More importantly, we show that LD converges to the optimal objective function values by the joint training under mild assumptions. Experiments demonstrate LD significantly outperforms state-of-the-art methods on Open Graph Benchmark datasets. Zhihao Shi, Jie Wang 0005, Fanghua Lu, Hanzhu Chen, Defu Lian, Jieping Ye, Feng Wu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Event-Based Stereo Depth Estimation by Temporal-Spatial Context LearningabstractEvent cameras represent a cutting-edge sensor technology, recording asynchronous pixel-level intensity changes with high temporal resolution and a wide dynamic range. These attributes make event-based stereo depth estimation particularly robust for scenarios characterized by rapid changes and challenging lighting conditions. However, previous learning-based approaches for event-based stereo have often overlooked exploiting the temporal context information within the scene, resulting in suboptimal depth estimations. In this paper, we introduce a novel learning-based network for event-based stereo that incorporates two innovative modules: the Event-based Temporal Aggregation Module (E-TAM) and the Temporal-guided Spatial Context Learning Module (T-SCLM). The E-TAM is designed to capture temporal context information among temporal features extracted from the entire event stream, further the T-SCLM exploits the temporal context information to provide guidance for spatial context learning. Subsequently, these merged features are input into the stereo matching network, ultimately yielding the final disparity map. Experimental evaluations conducted on two real-world datasets affirm the superiority of our method when compared to state-of-the-art approaches. Yueyi Zhang 0001, Xiaoyan Sun 0001, Feng Wu 0005 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Calculation of the Worst-Case Voltage Noise for a Power Distribution Network Based on Ramp CurrentabstractWith the continuous reduction of integrated circuit processing size and increasing integration density, power supply noise seriously threatens further improvement of high-speed digital system performance, which can cause jitter, latency, and even system malfunction. Power supply noise is strongly related to the design quality of the power distribution network and the activity of transient currents. In this paper, a method for estimating the worst-case voltage supply based on the ramp time of the transient current is proposed. At the same time, the worst-case current activity can also be obtained which can be used to guide the core or field programmable gate array (FPGA) design for a given power distribution network. The ramp current is first encoded and divided into segments. Then, two dynamic matrices are used to track the worst-case accumulated voltage noise and worst-case current activities. The proposed method is compared with state-of-the-art approaches showing consistent advantages in estimating the worst-case power supply noise. Its accuracy is also validated by comparing the simulated and measured data. Yuhuan Luo, Jun Wang 0034, Haiyue Yuan, Feng Wu 0005, Yang Liu 0091, Xiuqin Chu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | Graph-DETR4D: Spatio-Temporal Graph Modeling for Multi-View 3D Object DetectionabstractMulti-View 3D object detection (MV3D) has made tremendous progress by leveraging multiple perspective features through surrounding cameras. Despite demonstrating promising prospects in various applications, accurately detecting objects through camera view in the 3D space is extremely difficult due to the ill-posed issue in monocular depth estimation. Recently, Graph-DETR3D presents a novel graph-based 3D-2D query paradigm in aggregating multi-view images for 3D object detection and achieves competitive performance. Although it enriches the query representations with 2D image features through a learnable 3D graph, it still suffers from limited depth and velocity estimation abilities due to the adoption of a single-frame input setting. To solve this problem, we introduce a unified spatial-temporal graph modeling framework to fully leverage the multi-view imagery cues under the multi-frame inputs setting. Thanks to the flexibility and sparsity of the dynamic graph architecture, we lift the original 3D graph into the 4D space with an effective attention mechanism to automatically perceive imagery information at both spatial and temporal levels. Moreover, considering the main latency bottleneck lies in the image backbone, we propose a novel dense-sparse distillation framework for multi-view 3D object detection, to reduce the computational budget while sacrificing no detection accuracy, making it more suitable for real-world deployment. To this end, we propose Graph-DETR4D, a faster and stronger multi-view 3D object detection framework, built on top of Graph-DETR3D. Extensive experiments on nuScenes and Waymo benchmarks demonstrate the effectiveness and efficiency of Graph-DETR4D. Notably, our best model achieves 62.0% NDS on nuScenes test leaderboard. Code is available at https://github.com/zehuichen123/Graph-DETR4D. Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Wu 0005, Feng Zhao 0004 |
IEEE Trans. Image Process. | 7 |
| 2024 | Toward Decentralized Task Offloading and Resource Allocation in User-Centric MECabstractIn the traditional cellular-based mobile edge computing (MEC), users at the edge of the cell are prone to suffer severe inter-cell interference and signal attenuation, leading to low throughput even transmission interruptions. Such edge effect severely obstructs offloading of tasks to MEC servers. To address this issue, we propose user-centric mobile edge computing (UCMEC), a novel MEC architecture integrating user-centric transmission, which can ensure high throughput and reliable communication for task offloading. Then, we formulate an long-term delay minimization problem by jointly optimizing task offloading, power allocation, and computing resource allocation in UCMEC. To solve the intractable problem, we propose two decentralized joint optimization schemes based on multi-agent deep reinforcement learning (MADRL) and convex optimization, which consider both cooperation and non-cooperation among network nodes. Simulation results demonstrate that the proposed schemes in UCMEC can significantly improve the uplink transmission rate by at least 176.99% and reduce the long-term average total delay by at least 16.36% compared to traditional cellular-based MEC. Langtian Qin, Hancheng Lu, Baolin Chong, Feng Wu 0005 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Joint Optimization of Base Station Clustering and Service Caching in User-Centric MECabstractEdge service caching can effectively reduce the delay or bandwidth overhead for acquiring and initializing applications. To address single-base station (BS) transmission limitation and serious edge effect in traditional cellular-based edge service caching networks, in this paper, we proposed a novel user-centric edge service caching framework where each user is jointly provided with edge caching and wireless transmission services by a specific BS cluster instead of a single BS. To minimize the long-term average delay under the constraint of the caching cost, a mixed integer non-linear programming (MINLP) problem is formulated by jointly optimizing the BS clustering and service caching decisions. To tackle the problem, we propose JO-CDSD, an efficiently joint optimization algorithm based on Lyapunov optimization and generalized benders decomposition (GBD). In particular, the long-term optimization problem can be transformed into a primal problem and a master problem in each time slot that is much simpler to solve. The near-optimal clustering and caching strategy can be obtained through solving the primal and master problem alternately. Extensive simulations show that the proposed joint optimization algorithm outperforms other algorithms and can effectively reduce the long-term delay and caching cost. Langtian Qin, Hancheng Lu, Yao Lu 0024, Chenwu Zhang, Feng Wu 0005 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Vision-and-Language Navigation via Latent Semantic Alignment LearningabstractVision-and-Language Navigation (VLN) requires that an agent can comprehensively understand the given instructions and the immediate visual information obtained from the environment, so as to make correct actions to achieve the navigation goal. Therefore, semantic alignment across modalities is crucial for the agent understanding its own state during the navigation process. However, the potential of semantic alignment has not been systematically explored in current studies, which limits the further improvement of navigation performance. To address this issue, we propose a new Latent Semantic Alignment Learning method to develop the semantically aligned relationships contained in the environment. Specifically, we introduce three novel pre-training tasks: Trajectory-conditioned Masked Fragment Modeling, Action Prediction of Masked Observation, and Hierarchical Triple Contrastive Learning. The first two tasks are used to reason about cross-modal dependencies, while the third one is able to learn semantically consistent representations across modalities. In this way, the Latent Semantic Alignment Learning method establishes a consistent perception of the environment and makes the agent's actions easier to explain. Experiments on common benchmarks verify the effectiveness of our proposed methods. For example, we improve the Success Rate by 1.6% on the R2R validation unseen set and 4.3% on the R4R validation unseen set over the baseline model. Siying Wu, Xueyang Fu, Feng Wu 0005, Zhengjun Zha |
IEEE Trans. Multim. | 3 |
| 2023 | Fast Estimation of a Statistical Eye Diagram for Nonlinear High-Speed Links Based on the Minimum Required Order of the Multiple Edge Response MethodabstractWith the increase in nonlinear effects of high-speed links, the higher order multiple edge response (MER) method is widely used to accurately evaluate high-speed link performance by computing the statistical eye diagram at the price of massive edge responses. The computational complexity of the MER algorithm exponentially increases with the order. In this article, the estimation method for the minimum order of the MER method is proposed based on the last time of nonlinearity in high-speed links. The edge responses are divided into nonlinear and linear sections according to the minimum order of MER. To further improve the computational efficiency of the convolution operation of higher order MERs, a hybrid method is proposed that combines the double edge response (DER) method and statistical data of the edge response. The DER convolution is used to obtain the probability density distribution in the linear section, and the probability density distribution in the nonlinear section is obtained by the statistics of the rising and falling edge responses with different leading bits, which can greatly reduce the computation time. The accuracy and efficiency of the proposed method are verified by comparison with the higher order MER method. Jun Wang 0034, Yuhuan Luo, Wenting Guo, Feng Wu 0005, Xiuqin Chu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | Fast and Accurate Estimation of Statistical Eye Diagram for Nonlinear High-Speed LinksabstractA fast and accurate statistical eye diagram estimation method for high-speed nonlinear links is proposed in this article. Probability density functions (PDFs) of output responses are derived based on multiple edge responses (MERs). According to the property that the influence of nonlinearity will not propagate for a long time in high-speed links, a new scheme for calculating the PDFs of responses is presented, in which the convolution process is divided into nonlinear section, transition section, and linear section. Convolutions via high-order of MERs are only used for the nonlinear section, low-order of MERs are used for the transition section, and double edge responses are used for the linear section. The new scheme can drastically reduce the amount of computation. The proposed method is verified by comparing the probability density distributions of the statistical eye diagram, the bathtub curves, and the simulation time with that of the traditional total MER-based statistical eye diagram. Results show that the simulation speed of the proposed method has been improved by more than ten times, and the accuracy is almost the same as the traditional statistical eye diagram for nonlinear links. This method provides an efficient and accurate solution for estimating the statistical and BER eye diagrams for serious nonlinear links. Xiuqin Chu, Wenting Guo, Jun Wang 0034, Feng Wu 0005, Yuhuan Luo, Yushan Li 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |