Mingxuan Yu

dblp:352/8635 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and Encoding
abstract
Real-time multimodal inference on resource-constrained edge devices is essential for applications such as autonomous driving, human-computer interaction, and mobile health. However, prior work often overlooks the tight coupling between sensing dynamics and model execution, as well as the complex inter-modality dependencies. In this paper, we propose MMEdge, a new on-device multimodal inference framework based on pipelined sensing and encoding. Instead of waiting for complete sensor inputs, MMEdge decomposes the entire inference process into a sequence of fine-grained sensing and encoding units, allowing computation to proceed incrementally as data arrive. MMEdge also introduces a lightweight but effective temporal aggregation module that captures rich temporal dynamics across different pipelined units to maintain accuracy performance. Such pipelined design also opens up opportunities for fine-grained cross-modal optimization and early decision-making during inference. To further enhance system performance under resource variability and input data complexity, MMEdge incorporates an adaptive multimodal configuration optimizer that dynamically selects optimal sensing and model configurations for each modality under latency constraints, and a cross-modal speculative skipping mechanism that bypasses future units of slower modalities when early predictions reach sufficient confidence. We evaluate MMEdge using two public multimodal datasets and deploy it on a real-world unmanned aerial vehicle (UAV)-based multimodal testbed. The results show that MMEdge significantly reduces end-to-end latency while maintaining high task accuracy across various system and data dynamics. A video demonstration of MMEdge’s performance in real world is available at https://youtu.be/qRew7sT-iWw.
Runxi Huang, Mingxuan Yu, Mingyu Tsoi, Xiaomin Ouyang
SenSys2
2025 SeqBalance: Congestion-Aware Load Balancing With No Reordering in Data Center Networks
abstract
With the rapid development of the Internet of Things (IoT), an increasing amount of sensor data generated by IoT applications has been transferred to data center networks for storage and data analysis. Remote Direct Memory Access (RDMA) is widely used in data center networks because of its high performance. However, due to the characteristics of RDMA’s retransmission strategy, current load balancing schemes for data center networks are unsuitable for RDMA. In this paper, we propose SeqBalance, a load balancing framework designed for RDMA. SeqBalance implements fine-grained load balancing for RDMA through a reasonable design and does not cause reordering problems. SeqBalance detects link congestion at the switch by sensing ECN signals and link utilization, and guides routing decisions accordingly. SeqBalance’s designs are all based on existing commercial RNICs and commercial programmable switches, so they are compatible with existing data center networks. We have implemented SeqBalance Shaper for fine-grained sub-flow splitting in Mellanox CX-6 RNIC and implemented routing decisions in Intel Tofino P4 programmable switch. The results of hardware testbed experiments and large-scale simulations show that compared with existing load balancing schemes, SeqBalance improves 24.7% and 15.9% on average FCT and 99th-percentile FCT.
Huimin Luo, Jiao Zhang 0002, Mingxuan Yu, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
IEEE Internet Things J.3
2025 RoCELet: Host-Based Flowlet Load Balancing for RoCE
abstract
Remote Direct Memory Access (RDMA) is becoming a popular high-speed networking technology. It uses kernel bypass and zero copy to achieve high throughput and low latency with little CPU overhead. However, standard RoCE transmission uses Equal Cost Multipath (ECMP) for load balancing, which can result in lower transmission performance due to hash conflicts. Meanwhile, it has been verified that, unlike TCP, the unique retransmission mode and flow characteristics of RoCE make previous load balancing algorithms not well applied to RoCE. In this paper, we introduce RoCELet, a load balancing algorithm for RoCE. It achieves fine-grained RoCE load balancing by actively generating flowlets, effectively utilizing the rich end-to-end paths in the data center. We implement a prototype based on DPDK and evaluate it through small-scale testbed experiments and large-scale simulations. Our results show that compared to state-of-the-art load balancing algorithms, RoCELet optimizes 48.2% and 16.4% in average FCT and$99^{th}$-ile FCT, respectively.
Huimin Luo, Jiao Zhang 0002, Mingxuan Yu, Jiafeng Jiang, Yongchen Pan, Tian Pan 0001, Tao Huang 0005
IEEE Trans. Netw.3
2025 Re-Architecting Traffic Control in Cross-Datacenter RDMA Networks
abstract
The network-intensive applications, like machine learning and cloud storage, are increasingly driving two critical trends:1)RDMA has been widely deployed to provide high-speed networks;2)applications are distributively deployed across multiple regional datacenters to satisfy demands for content providers and customers. To fully utilize the benefits of RDMA, we desire to extend it to support cross-datacenter networks. However, the long-haul transport suffers a considerably long control loop, and thus the hybrid of long-haul and intra-datacenter traffic can easily cause severe congestion. We revisit existing traffic control methods and find they are insufficient to resolve this hybrid traffic congestion. Generally, regional datacenters are connected using dedicated long-haul optical fiber and datacenter interconnection (DCI) switches. In this paper, we propose Approach Traffic Control (ATC), a novel solution focusing on two-side DCI-switches (i.e., the approach point for datacenters) to separately alleviate the hybrid traffic congestion in the local and distal datacenters, as a building block for host-driven control methods. This design principle helps ATC shorten the control loop to a single datacenter scale while aggregating congestion information of the whole datacenter range with minor deployment complexity. We implement ATC on P4-based switches and conduct evaluations using real-world testbeds and large-scale NS3 simulations. The results show that ATC ensures fast congestion avoidance and delivers significant performance. For example, ATC reduces the FCT of intra-datacenter and long-haul traffic by up to 88% and 52%, respectively.
Zirui Wan, Jiao Zhang 0002, Yuzhen Su, Haoyu Pan, Mingxuan Yu, Tao Huang 0005
IEEE Trans. Netw.5
2024 BiCC: Bilateral Congestion Control in Cross-datacenter RDMA Networks
abstract
With the development of network-intensive applications like machine learning and cloud storage, there are two growing trends: (i) RDMA has been widely deployed to enhance underlying high-speed networks; (ii) applications are deployed on geographically distributed datacenters to meet customer demands (e.g., low access latency to services or regular data backups). To fully utilize the benefits of RDMA, we desire to support long-haul RDMA transport for cross-datacenter applications. Different from common intra-datacenter communications, the hybrid of long-haul and intra-datacenter traffic complicates the congestion state, and the considerably long control loop makes it more severe. We revisit existing congestion control methods and find they are insufficient to address the hybrid traffic congestion.Note that regional datacenters are connected by dedicated long-haul optical fiber and datacenter interconnection (DCI) switches directly. In this paper, we propose Bilateral Congestion Control (BiCC), a novel solution relying on two-side DCI-switches to bilaterally alleviate the hybrid traffic congestion in the sender-side and receiver-side datacenter while serving as a building block for existing host-driven methods. BiCC can shorten the control loop to a single datacenter scale and aggregate congestion information across the whole datacenter. We implement BiCC on commodity P4-based switches and conduct evaluations using both testbed experiments and NS3 simulations. The extensive evaluation results show that BiCC ensures fast congestion avoidance. Thus, BiCC reduces the average FCT for intra-datacenter and inter-datacenter traffic by up to 53% and 51%, respectively, in large-scale simulations.
Zirui Wan, Jiao Zhang 0002, Mingxuan Yu, Xinghua Zhao, Tao Huang 0005
INFOCOM3
2024 PACC: A Proactive CNP Generation Scheme for Datacenter Networks
abstract
The rapid upgrade of link speed and the prosperity of new applications in data center networks (DCNs) lead to a rigorous demand for ultra-low latency and high throughput. To mitigate the overhead of traditional software-based packet processing at end-hosts, RDMA (Remote Direct Memory Access) has been widely adopted in DCNs. Particularly, congestion control (CC) mechanisms designed for RDMA have attracted much attention to avoid performance deterioration when packets lose. However, through comprehensive analysis, we found that existing RDMA CC schemes have limitations of a sluggish response to congestion and unawareness of tiny microbursts due to the long end-to-end control loop. In this paper, we propose PACC, a proactive and accurate switch-driven RDMA CC algorithm with easy deployability. PACC is driven by PI controller-based computation, threshold-based flow discrimination and weight-based allocation at the switch. It leverages real-time queue length to generate accurate congestion feedback proactively and piggybacks it to the corresponding source without modification to end-hosts. We theoretically analyze the stability, convergence and key parameter settings of PACC. Then, we implement PACC in a testbed consisting of DPDK-based end-hosts and Tofino P4 switches. In our evaluation, PACC achieves better fairness, fast reaction, high throughput, and 6$\sim$69% lower FCT (Flow Completion Time) than DCQCN, TIMELY, HPCC and RoCC.
Jiao Zhang 0002, Xiaolong Zhong, Mingxuan Yu, Haoyu Pan, Zixuan Guan, Biyao Che, Zirui Wan, Tian Pan 0001, Tao Huang 0005
IEEE/ACM Trans. Netw.4
2023 Gradient Corner Pooling for Keypoint-Based Object Detection
abstract
Detecting objects as multiple keypoints is an important approach in the anchor-free object detection methods while corner pooling is an effective feature encoding method for corner positioning. The corners of the bounding box are located by summing the feature maps which are max-pooled in the x and y directions respectively by corner pooling. In the unidirectional max pooling operation, the features of the densely arranged objects of the same class are prone to occlusion. To this end, we propose a method named Gradient Corner Pooling. The spatial distance information of objects on the feature map is encoded during the unidirectional pooling process, which effectively alleviates the occlusion of the homogeneous object features. Further, the computational complexity of gradient corner pooling is the same as traditional corner pooling and hence it can be implemented efficiently. Gradient corner pooling obtains consistent improvements for various keypoint-based methods by directly replacing corner pooling. We verify the gradient corner pooling algorithm on the dataset and in real scenarios, respectively. The networks with gradient corner pooling located the corner points earlier in the training process and achieve an average accuracy improvement of 0.2%-1.6% on the MS-COCO dataset. The detectors with gradient corner pooling show better angle adaptability for arrayed objects in the actual scene test.
Xuemei Xie, Mingxuan Yu, Jiakai Luo, Chengwei Rao, Guangming Shi
AAAI3
2023 OCSKB: An Object Component Sketch Knowledge Base for Fast 6D Pose Estimation
abstract
6D pose estimation from a single RGB image is a fundamental task in computer vision. In most methods of instance-level or category-level 6D pose estimation, accurate CAD models or point cloud models are indispensable part. It is not easy to quickly obtain the models of these everyday objects. To address this issue, we present a part-level object component sketch knowledge base which consists of 270 real-world object sketch models of 30 categories. Objects are disassembled into geometry components with spatial relationship according to their functions and structures, and convert them into three basic spatial structures: frustum, circular truncated cone, and sphere. We present a fast pipeline for sketch modeling with our tool. The average time for this method to build a simple model for everyday objects is about 2 minutes. Additionally, we leverage the geometric information and spatial relationships inherent in the multiple viewpoint projection maps of these sketch bases to develop a rapid inference framework for 6D pose estimation. The interpretable steps in our framework gradually retrieve and activate valid solutions in the discrete 6D pose space. Extensive experiments in real-world environments have demonstrated that our method can reliably and robustly estimate the 6D pose of objects, even without access to accurate CAD or point cloud models. Furthermore, our method achieves state-of-the-art performance, operating at a speed of 90 frames per second using parallel computing on GPU.
Guangming Shi, Xuemei Xie, Mingxuan Yu, Chengwei Rao, Jiakai Luo
ACM Multimedia4