Weijian Xu

dblp:55/5787 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 6 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Small-World Topology and Graph Attention Reinforcement Learning for dynamic traffic optimization
Zhibin Gao, Zhongzhe Song, Yanglong Sun, Weijian Xu, Lianyou Lai, Shuwu Chen
Eng. Appl. Artif. Intell.5
2026 FPGA-Based Adjoint Method Accelerator for Rapid Optical Inverse Design
abstract
Inverse design is an emerging theory in material, photonics, and other fields. The Adjoint Method (AM) is an efficient gradient-based optimization strategy commonly used in inverse design. The optimization of inverse design frameworks faces fundamental challenges in achieving computational efficiency and scalability. While previous approaches have explored both algorithmic optimization and hardware acceleration, the inherent trade-offs between memory consumption and processing throughput continue to limit performance at scale. In this paper, we propose a high-speed field-programmable gate array (FPGA)-based accelerator to enhance the performance of the AM for fast inverse design. The accelerator is designed to significantly reduce computation time and enhance scalability, enabling more efficient and rapid design iterations. The accelerator is implemented on an AMD Alveo U280 FPGA, and we compared the FPGA and graphic process unit (GPU) performance on a series of inverse design tasks. The experimental results demonstrate that the proposed FPGA accelerator offers superior performance in terms of speed, especially in smaller design sizes compared to the NVIDIA A100 GPU. This significant improvement is due, in part, to the fully pipelined architecture and the low memory bandwidth requirement of the accelerator, which efficiently handles the enormous data throughput required for these tasks.
Lianyou Lai, Zhiheng Ni, Miaoxiang Yu, Zhenyu Xu 0007, Weijian Xu
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 Olympus: A Universal Task Router for Computer Vision Tasks
abstract
We introduce Olympus, a new approach that transforms Multimodal Large Language Models (MLLMs) into a unified framework capable of handling a wide array of computer vision tasks. Utilizing a controller MLLM, Olympus delegates over 20 specialized tasks across images, videos, and 3D objects to dedicated modules. This instruction-based routing enables complex workflows through chained actions without the need for training heavy generative models. Olympus easily integrates with existing MLLMs, expanding their capabilities with comparable performance. Experimental results demonstrate that Olympus achieves an average routing accuracy of 94.75% across 20 tasks and precision of 91.82% in chained action scenarios, showcasing its effectiveness as a universal task router that can solve a diverse range of computer vision tasks.
Yuanze Lin, Yunsheng Li, Dongdong Chen 0001, Weijian Xu, Ronald Clark, Philip Torr 0001
CVPR4
2025 MADRL-Based Edge Computing: Joint Energy-Latency Optimization for Marine Internet of Things
abstract
Mobile edge computing technology has facilitated the deployment of high computational algorithms on Marine Internet of things (MIoT) devices that equipped with limited computing resources. However, the dynamic changes in the network environment, the strict requirements of system latency and energy consumption of mobile devices, restrict the task execution in MIoT. This paper proposes an offloading framework in a marine mobile edge computing scenario, which is assisted by sea buoys and unmanned aerial vehicles (UAVs). Aiming to achieve long-term system optimization goals under resource-constrained conditions, with task partitioning, user scheduling, and resource allocation joint optimization under the constraints of UAV residual battery life, system latency, and energy consumption. This paper specifically introduces a Network Partition Point Reservation Algorithm (NPPR) that reduces the solution space of the problem through preprocessing. Subsequently, the paper presents a multiagent deep deterministic policy gradient (MADDPG) algorithm, enhanced with standardization and adaptive learning rate decay (SA-MADDPG), to address task heterogeneity and environmental dynamics. Simulation results demonstrate that, compared to existing algorithms, the SA-MADDPG algorithm proposed in this paper reduces system latency and energy consumption by 34.7% and 61.2%, respectively.
Weijian Xu, Wenqian Luo, Yanglong Sun, Zhibin Gao, Lianyou Lai
IEEE Internet Things J.1
2024 Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
abstract
We introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for various computer vision and vision-language tasks. While existing large vision models excel in transfer learning, they struggle to perform diverse tasks with simple instructions, a capability that implies handling the complexity of various spatial hierarchy and semantic granularity. Florence-2 was designed to take text-prompt as task instructions and generate desirable results in text forms, whether it be captioning, object detection, grounding or segmentation. This multi-task learning setup demands large-scale, high-quality annotated data. To this end, we co-developed FLD-5B that consists of 5.4 billion comprehensive visual annotations on 126 million images, using an iterative strategy of automated image annotation and model refinement. We adopted a sequence-to-sequence structure to train Florence-2 to perform versatile and comprehensive vision tasks. Extensive evaluations on numerous tasks demonstrated Florence-2 to be a strong vision foundation model contender with un-precedented zero-shot and fine-tuning capabilities.
Bin Xiao 0004, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng 0001, Ce Liu 0001, Lu Yuan 0001
CVPR3
2024 Dynamic Spatiotemporal Graph Wavelet Network for Traffic Flow Prediction
abstract
Real-time and high-precision traffic flow prediction plays a crucial role in transportation management, contributing to control dispatch and reducing traffic congestion. Due to the complex and dynamic characteristics of the traffic flow, traffic flow prediction remains challenging. Previous work often ignores some dynamic and momentary spatial information, and predicting the traffic flow of a target road segment in the long-term horizon is a difficult problem. To address these issues, we propose a novel deep learning framework, termed long short-term structural spatiotemporal information fusion graph wavelet network (LSSTF-GWN), to capture momentary dynamic spatiotemporal correlation and make long-term predictions. The LSSTF-GWN model integrates graph wavelets network (GWN) with temporal gated convolution networks into graph convolution network to construct a multigraph architecture to address the complex spatiotemporal correlations in traffic flow data. The GWN extracts the instantaneous and global features of the spatial information by designing different adjacency matrices. The LSSTF-GWN not only considers the fixed distance between graphs but also builds long-term dynamic graphs for the inner relationships of the nodes to reflect contextual and global information. The proposed method is evaluated using three metrics, i.e., mean absolute error, root-mean-square error, and mean absolute percentage error, on two real-world data sets from the Caltrans Performance Measurement System (PeMS). The experimental results demonstrate the superior performance of our method in long-term traffic flow prediction.
Weijian Xu, Jingjin Liu, Juan Yang 0005, Huifen Liu, Teng Zhou
IEEE Internet Things J.1
2024 Latency-Aware MIoT Service Strategy in UAV-Assisted Dynamic MMEC Environment
abstract
Marine Internet of Things (MIoT) has emerged as a prominent technology for the future development of marine applications, in which edge equipment provides a valuable method for information collection and processing on smart mobile devices (SMDs). However, the deployment of edge equipment may result in high latency due to inefficient computing offloading schemes. In this paper, we propose an optimal offloading scheme based on a dynamic unmanned air vehicle (UAV) assisted marine mobile edge computing (MMEC) environment in which latency-sensitive computing tasks can be partially offloaded autonomously. Specifically, we consider a time-varying scenario where the UAV hovers over multiple maritime mobile unmanned surface vessels (USVs) and provides MIoT services over communication periods. Our objective is to minimize the overall task execution time through joint optimization of user scheduling variables, UAV motion trajectory, and resource allocation while considering energy consumption and spatial constraints, thereby achieving enhanced quality of service. Considering the non-convexity of this optimization problem, we propose an advanced Twin Delayed Deep Deterministic policy gradient (ATD3) algorithm and examine the convergence and optimality of different parameter factors. Simulation results demonstrate that the proposed algorithm is superior to the baseline scheme regarding convergence speed, adaptability, and task execution time.
Weijian Xu, Zhongzhe Song, Zhibin Gao, Lianyou Lai, Yanglong Sun, Wenqian Luo
IEEE Internet Things J.1
2022 Instance Segmentation with Mask-supervised Polygonal Boundary Transformers
abstract
In this paper, we present an end-to-end instance segmentation method that regresses a polygonal boundary for each object instance. This sparse, vectorized boundary representation for objects, while attractive in many downstream computer vision tasks, quickly runs into issues of parity that need to be addressed: parity in supervision and parity in performance when compared to existing pixel-based methods. This is due in part to object instances being annotated with ground-truth in the form of polygonal boundaries or segmentation masks, yet being evaluated in a convenient manner using only segmentation masks. Our method, BoundaryFormer, is a Transformer based architecture that directly predicts polygons yet uses instance mask segmentations as the ground-truth supervision for computing the loss. We achieve this by developing an end-to-end differentiable model that solely relies on supervision within the mask space through differentiable rasterization. Boundary-Former matches or surpasses the Mask R-CNN method in terms of instance segmentation quality on both COCO and Cityscapes while exhibiting significantly better transferability across datasets.
Justin Lazarow, Weijian Xu, Zhuowen Tu
CVPR2
2021 Convolutions and Self-Attention: Re-interpreting Relative Positions in Pre-trained Language Models
abstract
Tyler Chang, Yifan Xu, Weijian Xu, Zhuowen Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Tyler A. Chang, Yifan Xu 0009, Weijian Xu, Zhuowen Tu
ACL/IJCNLP (1)3
2021 Pose Recognition With Cascade Transformers
abstract
In this paper, we present a regression-based pose recognition method using cascade Transformers. One way to categorize the existing approaches in this domain is to separate them into 1). heatmap-based and 2). regression-based. In general, heatmap-based methods achieve higher accuracy but are subject to various heuristic designs (not end-to-end mostly), whereas regression-based approaches attain relatively lower accuracy but they have less intermediate non-differentiable steps. Here we utilize the encoder-decoder structure in Transformers to perform regression-based person and keypoint detection that is general-purpose and requires less heuristic design compared with the existing approaches. We demonstrate the keypoint hypothesis (query) refinement process across different self-attention layers to reveal the recursive self-attention mechanism in Transformers. In the experiments, we report competitive results for pose recognition when compared with the competing regression-based methods.
Kenneth Li 0002, Xiang Zhang 0015, Yifan Xu 0009, Weijian Xu, Zhuowen Tu
CVPR5
2021 Line Segment Detection Using Transformers Without Edges
abstract
In this paper, we present a joint end-to-end line segment detection algorithm using Transformers that is post-processing and heuristics-guided intermediate processing (edge/junction/region detection) free. Our method, named LinE segment TRansformers (LETR), takes advantages of having integrated tokenized queries, a self-attention mechanism, and encoding-decoding strategy within Transformers by skipping standard heuristic designs for the edge element detection and perceptual grouping processes. We equip Transformers with a multi-scale encoder/decoder strategy to perform fine-grained line segment detection under a direct endpoint distance loss. This loss term is particularly suitable for detecting geometric structures such as line segments that are not conveniently represented by the standard bounding box representations. The Transformers learn to gradually refine line segments through layers of self-attention. In our experiments, we show state-of-the-art results on Wireframe and YorkUrban benchmarks.
Yifan Xu 0009, Weijian Xu, David Cheung, Zhuowen Tu
CVPR2
2021 Co-Scale Conv-Attentional Image Transformers
abstract
In this paper, we present Co-scale conv-attentional image Transformers (CoaT), a Transformer-based image classifier equipped with co-scale and conv-attentional mechanisms. First, the co-scale mechanism maintains the integrity of Transformers’ encoder branches at individual scales, while allowing representations learned at different scales to effectively communicate with each other; we design a series of serial and parallel blocks to realize the co-scale mechanism. Second, we devise a conv-attentional mechanism by realizing a relative position embedding formulation in the factorized attention module with an efficient convolution-like implementation. CoaT empowers image Transformers with enriched multi-scale and contextual modeling capabilities. On ImageNet, relatively small CoaT models attain superior classification results compared with similar-sized convolutional neural networks and image/vision Transformers. The effectiveness of CoaT’s backbone is also illustrated on object detection and instance segmentation, demonstrating its applicability to downstream computer vision tasks.
Weijian Xu, Yifan Xu 0009, Tyler A. Chang, Zhuowen Tu
ICCV1
2021 Attentional Constellation Nets for Few-Shot Learning
Weijian Xu, Yifan Xu 0009, Huaijin Wang 0002, Zhuowen Tu
ICLR1
2020 Guided Variational Autoencoder for Disentanglement Learning
abstract
We propose an algorithm, guided variational autoencoder (Guided-VAE), that is able to learn a controllable generative model by performing latent representation disentanglement learning. The learning objective is achieved by providing signal to the latent encoding/embedding in VAE without changing its main backbone architecture, hence retaining the desirable properties of the VAE. We design an unsupervised and a supervised strategy in Guided-VAE and observe enhanced modeling and controlling capability over the vanilla VAE. In the unsupervised strategy, we guide the VAE learning by introducing a lightweight decoder that learns latent geometric transformation and principal components; in the supervised strategy, we use an adversarial excitation and inhibition mechanism to encourage the disentanglement of the latent variables. Guided-VAE enjoys its transparency and simplicity for the general representation learning task, as well as disentanglement learning. On a number of experiments for representation learning, improved synthesis/sampling, better disentanglement for classification, and reduced classification errors in meta learning have been observed.
Zheng Ding, Yifan Xu 0009, Weijian Xu, Gaurav Parmar, Yang Yang 0010, Max Welling, Zhuowen Tu
CVPR3
2019 3D Volumetric Modeling with Introspective Neural Networks
abstract
In this paper, we study the 3D volumetric modeling problem by adopting the Wasserstein introspective neural networks method (WINN) that was previously applied to 2D static images. We name our algorithm 3DWINN which enjoys the same properties as WINN in the 2D case: being simultaneously generative and discriminative. Compared to the existing 3D volumetric modeling approaches, 3DWINN demonstrates competitive results on several benchmarks in both the generation and the classification tasks. In addition to the standard inception score, the Frechet Inception Distance (FID) metric is´ also adopted to measure the quality of 3D volumetric generations. In addition, we study adversarial attacks for volumetric data and demonstrate the robustness of 3DWINN against adversarial examples while achieving appealing results in both classification and generation within a single model. 3DWINN is a general framework and it can be applied to the emerging tasks for 3D object and scene modeling.1
Wenlong Huang, Brian Lai, Weijian Xu, Zhuowen Tu
AAAI3
2019 Geometry-Aware End-to-End Skeleton Detection
Weijian Xu, Gaurav Parmar, Zhuowen Tu
BMVC1
2019 Optimal Joint Offloading and Wireless Scheduling for Parallel Computing with Deadlines
abstract
In this paper, we consider the problem of joint offloading and wireless scheduling design for parallel computing applications with hard deadlines. This is motivated by the rapid growth of compute-intensive mobile parallel computing applications (e.g., real-time video analysis, language translation) that require to be processed within a hard deadline. While there are many works on joint computing and communication algorithm design, most of them focused on the minimization of average computing time and may not be applicable for mobile applications with hard deadlines. In this work, we explicitly take hard deadlines for computing tasks into account and develop a joint offloading and scheduling algorithm based on the stochastic network optimization framework. The proposed algorithm is shown to achieve average energy consumption arbitrarily close to the optimal one. However, this algorithm involves a strong coupling between offloading and scheduling decisions, which yields significant challenges on its implementation. Towards this end, we first successfully decouple the offloading and scheduling decisions in the case with one time slot deadline by exploring the intrinsic structure of the proposed algorithm. Based on this, we further implement the proposed algorithm in the general setups. Simulations are provided to corroborate our findings.
Xudong Qin, Weijian Xu, Bin Li 0014
WiOpt2
2018 Wasserstein Introspective Neural Networks
abstract
We present Wasserstein introspective neural networks (WINN) that are both a generator and a discriminator within a single model. WINN provides a significant improvement over the recent introspective neural networks (INN) method by enhancing INN's generative modeling capability. WINN has three interesting properties: (1) A mathematical connection between the formulation of the INN algorithm and that of Wasserstein generative adversarial networks (WGAN) is made. (2) The explicit adoption of the Wasserstein distance into INN results in a large enhancement to INN, achieving compelling results even with a single classifier - e.g., providing nearly a 20 times reduction in model size over INN for unsupervised generative modeling. (3) When applied to supervised classification, WINN also gives rise to improved robustness against adversarial examples in terms of the error reduction. In the experiments, we report encouraging results on unsupervised learning problems including texture, face, and object modeling, as well as a supervised classification task against adversarial attacks. Our code is available online1.
Kwonjoon Lee, Weijian Xu, Fan Fan 0001, Zhuowen Tu
CVPR2
2017 Hierarchical content importance-based video quality assessment for HEVC encoded videos transmitted over LTE networks
Jiefeng Guo, Gong Hu, Weijian Xu, Lianfen Huang
J. Vis. Commun. Image Represent.3
2015 The RR-PEVQ algorithm research based on active area detection for big data applications
Weijian Xu, Caidan Zhao, Hua-Pei Chiang, Lianfen Huang, Yueh-Min Huang
Multim. Tools Appl.1