Tuo Cao

dblp:296/4868 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
22since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 iG-6DoF: Model-free 6DoF Pose Estimation for Unseen Object via Iterative 3D Gaussian Splatting
abstract
Traditional methods in pose estimation often rely on precise 3D models or additional data such as depth and normals, limiting their generalization, especially when objects undergo large translations or rotations. We propose iG6DoF, a novel model-free 6D pose estimation method using iterative 3D Gaussian Splatting to estimate the pose of unseen objects. We first estimates an initial pose by leveraging multi-scale data augmentation and the rotation-equivariant features to create a better pose hypothesis from a set of candidates. Then, we propose an iterative 3DGS approach through iteratively rendering and comparing the rendered image with the input image to further progressively improve pose estimation accuracy. The proposed method consists of an object detector, a multi-scale rotation-equivariant feature based initial pose estimator, and a coarse-to-fine pose refiner. Such combination allows our method to focus on the target object in a complex scene dealing with large movement and weak textures. Our method achieves state-of-the-art results on the LINEMOD, OnePose-LowTexture, GenMOP datasets and our self-captured data, demonstrating its strong generalization to unseen objects and robustness across various scenes.
Tuo Cao, Fei Luo 0004, Jiongming Qin, Yu Jiang 0007, Yusen Wang 0002, Chunxia Xiao
CVPR1
2025 Toward auction-based edge AI: Orchestrating and incentivizing online transfer learning in edge networks
Yang Chen 0001, Lei Jiao 0002, Tuo Cao, Ji Qi 0005, Gangyi Luo, Sheng Zhang 0001, Sanglu Lu, Zhuzhong Qian
Comput. Networks3
2025 Towards strong continuous consistency in edge-assisted VR-SGs: Delay-differences sensitive online task redistribution
Yunqi Sun, Hesheng Sun, Tuo Cao, Mingtao Ji, Zhuzhong Qian, Lingkun Meng
Comput. Networks3
2025 Orchestrating In-Network Aggregation for Distributed Machine Learning via In-Band Network Telemetry
Mingtao Ji, Yibo Jin 0001, Zhuzhong Qian, Tuo Cao
J. Comput. Sci. Technol.4
2024 Low-Carbon Geographically Distributed Cloud-Edge Task Scheduling
Yingjie Zhu, Ji Qi 0005, Shengjie Wei, Tuo Cao, Gangyi Luo, Zhuzhong Qian
ICA3PP (6)6
2024 HS-Surf: A Novel High-Frequency Surface Shell Radiance Field to Improve Large-Scale Scene Rendering
abstract
Previous neural radiance fields often struggle to preserve high-frequency textures in urban and aerial large-scale scenes due to insufficient model capacity on the scene surface. This is attributed to their sampling locations or grid vertices falling in empty areas. Additionally, most models do not consider the drastic changes in distances. To address these issues, we propose a novel high-frequency surface shell radiance field, which uses depth-guided information to create a shell enveloping the scene surface under the current view, and then samples conic frustums on this shell to render high-frequency textures. Specifically, our method comprises three parts. Initially, we propose a strategy to fuse voxel grids and information of distance scales to generate a coarse scene at different distance scales. Subsequently, we construct a shell based on the depth information to carry out compensation to incorporate texture details not captured by voxels. Finally, the smooth and denoise post-processing further improves the rendering quality. Substantial scene experiments and ablation experiments demonstrate that our method achieves the obvious improvement of high-frequency textures at different distance scales and outperforms the state-of-the-art methods.
Jiongming Qin, Fei Luo 0004, Tuo Cao, Wenju Xu, Chunxia Xiao
ACM Multimedia3
2024 Walking on two legs: Joint service placement and computation configuration for provisioning containerized services at edges
Tuo Cao, Qinhui Wang, Zhuzhong Qian, Yue Zeng 0002, Mingtao Ji, Hesheng Sun
Comput. Networks1
2024 INTaaS: Provisioning In-band Network Telemetry as a service via online learning
Mingtao Ji, Chenwei Su, Yitao Fan, Yibo Jin 0001, Zhuzhong Qian, Yu Chen 0038, Tuo Cao, Sheng Zhang 0001
Comput. Networks8
2024 DGECN++: A Depth-Guided Edge Convolutional Network for End-to-End 6D Pose Estimation via Attention Mechanism
abstract
Monocular object 6D pose estimation is a fundamental yet challenging task in computer vision. Recently, deep learning has been proven to be capable of predicting remarkable results in this task. Existing works often adopt a two-stage pipeline with establishing 2D-3D correspondences and utilizing a PnP/RANSAC or differentiable PnP algorithm to recover 6 degrees-of-freedom (6DoF) pose parameters. However, most of them hardly consider the geometric features in 3D space, and ignore the topological cues when performing differentiable PnP algorithms. To this end, we present an improved end-to-end monocular 6D pose estimation method (DGECN++) that incorporates depth estimation and a geometric-aware learnable PnP network. Our method is based on keypoints. First we detect the 2D keypoints that correspond to the 3D model. We then integrate differentiable PnP/RANSAC algorithm to create an end-to-end pipeline for 6D pose estimation. We focuses on the following three key aspects: 1) We utilize the estimated depth information to guide the process of extracting 2D-3D correspondences and refine the results using a cascaded differentiable PnP/RANSAC algorithm that incorporates geometric information. 2) We leverage the uncertainty of the estimated depth map to enhance the accuracy and robustness of the predicted 6D pose. 3) We propose a differentiable Perspective-n-Point (PnP) algorithm based on edge convolution and self-attention to explore the topological relationships between 2D-3D correspondences. Experimental results demonstrate that our proposed network surpasses existing methods in terms of both effectiveness and efficiency.
Tuo Cao, Yanping Fu, Shengjie Zheng, Fei Luo 0004, Chunxia Xiao
IEEE Trans. Circuits Syst. Video Technol.1
2023 Joint Video Transcoding and Representation Selection for Edge-Assisted Multi-party Video Conferencing
Fanhao Kong, Tuo Cao, Zhuzhong Qian, Xiaoliang Wang 0001, Zhenjie Lin
ICA3PP (1)2
2023 NeTO: Neural Reconstruction of Transparent Objects with Self-Occlusion Aware Refraction-Tracing
abstract
We present a novel method called NeTO, for capturing the 3D geometry of solid transparent objects from 2D images via volume rendering. Reconstructing transparent objects is a very challenging task, which is ill-suited for general-purpose reconstruction techniques due to the specular light transport phenomena. Although existing refraction-tracing-based methods, designed especially for this task, achieve impressive results, they still suffer from unstable optimization and loss of fine details since the explicit surface representation they adopted is difficult to be optimized, and the self-occlusion problem is ignored for refraction-tracing. In this paper, we propose to leverage implicit Signed Distance Function (SDF) as surface representation and optimize the SDF field via volume rendering with a self-occlusion aware refractive ray tracing. The implicit representation enables our method to be capable of reconstructing high-quality reconstruction even with a limited set of views, and the self-occlusion aware strategy makes it possible for our method to accurately reconstruct the self-occluded regions. Experiments show that our method achieves faithful reconstruction results and outperforms prior works by a large margin. Visit our project page at https://www.xxlong.site/NeTO/.
Zongcheng Li, Xiaoxiao Long, Yusen Wang 0002, Tuo Cao, Wenping Wang 0001, Fei Luo 0004, Chunxia Xiao
ICCV4
2023 BIRP: Batch-aware Inference Workload Redistribution and Parallel Scheme for Edge Collaboration
abstract
The inference workload redistribution is a technique for evacuating inference requests from hot edges to idle edges in edge collaborative systems, thereby achieving inference workload balancing for inference on different edges. However, with the continuous development of edge accelerators, the resource utilization of edge accelerators in executing inference requests in series is often low, and when executing multiple inference requests in parallel, it faces uncertain execution delays, different response-time Service Level Objectives (SLOs), and the generality of inference workloads in heterogeneous edge collaborative systems. To address these issues, for the first time in the domain of inference workload redistribution, we propose a Batch-aware Inference workload Redistribution and Parallel execution scheme, called BIRP, to reduce the additional latency caused by waiting for a single inference task during serial execution, thereby improving the overall inference accuracy. BIRP uses the Multi-Armed Bandit (MAB) algorithm to adjust hyperparameters of the Throughput Improvement Ratio (TIR) function online for improving the overall inference accuracy. For nonlinear terms in the problem, BIRP uses a piecewise linear approximation to convert it into a Quadratic Programming (QP) problem, ensuring the effectiveness of BIRP in theory. We prototype BIRP on an edge collaborative system composed of three heterogeneous edges. Based on real inference workload trace, we validate the superiority of our algorithm compared to the state-of-the-art model selection-based inference workload redistribution algorithm, with an overall inference loss reduction of at least 32.9% and the failure rate of SLO has been reduced to 19.8% of alternatives.
Hesheng Sun, Zhuzhong Qian, Zengji Li, Ning Chen 0010, Tuo Cao, Suwei Xu
ICPP6
2023 Adaptive Provisioning In-band Network Telemetry at Computing Power Network [invited]
abstract
In-band Network Telemetry (INT) is proposed to detect networks via injecting specific probes to collect the hop-by-hop metadata within programmable switches. But there exist multiple challenges to conducting INT at Computing Power Network, such as control decisions of different INT frequencies, and the unforeseeable INT query workloads. In this study, we formulate an online non-linear time-varying integer programming problem that aims to maximize the overall quality of service through both frequency selection and INT query workload distribution. To achieve this, we propose an online learning, INTService, which utilizes a primal-dual mechanism to make fractional decisions. At last, extensive evaluations show that our proposed INTService exhibits up-lift performance 40% on average over other state-of-the-art algorithms.
Mingtao Ji, Chenwei Su, Zhuzhong Qian, Sheng Zhang 0001, Yu Chen 0038, Tuo Cao, Xiaohang Shi 0001, Luis Vasquez
IWQoS7
2023 Incentivizing Edge AI with Accuracy Preserving via Online Randomized Auctions
abstract
Provisioning machine learning inference near the users at the network’s edge is emerging as a promising area for Edge AI. Due to the excessive energy consumption of edge devices, their owners often lack the motivation to actively contribute to their edge resources. To tackle this issue, we propose an incentive mechanism based on auctions, which enables edge device owners to submit their bids and compensates such bids via reward payments, ensuring the minimization of inference accuracy loss. We formulate a nonlinear mixed-integer program problem with the objective of minimizing the social cost, including accuracy loss cost, edge device cost, and service provider cost in the edge inference system. Then an Online Learning algorithm is devised to find the solutions, based on the primal-dual. To calculate the remuneration, we design a payment allocation algorithm based on the bid-winning probabilities. Our rigorous theoretical analysis shows that our algorithm designed achieves sub-linear growth on dynamic regret and dynamic fit over time while preserving the economic properties of truthfulness and individual rationality. Finally, multiple experiments validate the efficacy of the proposed auction mechanism algorithm from various perspectives compared with three other existing algorithms.
Mingtao Ji, Zhuzhong Qian, Tuo Cao, Chenwei Su
SECON5
2023 Self-Supervised Monocular Depth Estimation by Digging into Uncertainty Quantification
Yuanzhen Li, Shengjie Zheng, Zi-Xin Tan, Tuo Cao, Fei Luo 0004, Chunxia Xiao
J. Comput. Sci. Technol.4
2022 DGECN: A Depth-Guided Edge Convolutional Network for End-to-End 6D Pose Estimation
abstract
Monocular 6D pose estimation is a fundamental task in computer vision. Existing works often adopt a two-stage pipeline by establishing correspondences and utilizing a RANSAC algorithm to calculate 6 degrees-of-freedom (6DoF) pose. Recent works try to integrate differentiable RANSAC algorithms to achieve an end-to-end 6D pose estimation. However, most of them hardly consider the geometric features in 3D space, and ignore the topology cues when performing differentiable RANSAC algorithms. To this end, we proposed a Depth-Guided Edge Convolutional Network (DGECN) for 6D pose estimation task. We have made efforts from the following three aspects: 1) We take advantages of estimated depth information to guide both the correspondences-extraction process and the cascaded differentiable RANSAC algorithm with geometric information. 2) We leverage the uncertainty of the estimated depth map to improve accuracy and robustness of the output 6D pose. 3) We propose a differentiable Perspective-n-Point(PnP) algorithm via edge convolution to explore the topology relations between 2D-3D correspondences. Experiments demonstrate that our proposed network outperforms current works on both effectiveness and efficiency.
Tuo Cao, Fei Luo 0004, Yanping Fu, Shengjie Zheng, Chunxia Xiao
CVPR1
2022 Towards Energy-efficient Container Data Center: An Online Migratability-aware Orchestrator
abstract
The growing demands for cloud computing have led to high power consumption in data centers. To save power, existing works attempt to reduce the number of active servers via scheduling or migrating VMs. Meanwhile, owing to the lightweight, highly-portable and scalable properties, containers have been widely used in data centers. However, comparing with VMs, containers show the migratability property, i.e., some containers can not be consolidated by migration. Thus, one has to take this property into consideration when making scheduling and migrating decisions. To this end, this paper studies the energy-efficient container orchestration problem in data centers. We frame container scheduling and migration as a single problem, and take the migratability of containers into account. We propose an online container orchestration algorithm, based on Lyapunov optimization and Markov approximation. It works without requiring future information and achieves a provable performance guarantee. Simulation results show that our algorithm can effectively reduce the power consumption of data centers.
Shengjie Wei, Tuo Cao, Sheng Zhang 0001, Zhuzhong Qian
MSN3
2022 Adaptive provisioning for mobile cloud gaming at edges
Tuo Cao, Yibo Jin 0001, Xiongfeng Hu, Sheng Zhang 0001, Zhuzhong Qian, Sanglu Lu
Comput. Networks1
2022 Deep attentive style transfer for images with wavelet decomposition
Gang Fu 0003, Qinan Yan, Caoqing Jiang, Tuo Cao, Shenghong Hu, Chunxia Xiao
Inf. Sci.5
2022 NeuralRoom: Geometry-Constrained Neural Implicit Surfaces for Indoor Scene Reconstruction
abstract
We present a novel neural surface reconstruction method called NeuralRoom for reconstructing room-sized indoor scenes directly from a set of 2D images. Recently, implicit neural representations have become a promising way to reconstruct surfaces from multiview images due to their high-quality results and simplicity. However, implicit neural representations usually cannot reconstruct indoor scenes well because they suffer severe shape-radiance ambiguity. We assume that the indoor scene consists of texture-rich and flat texture-less regions. In texture-rich regions, the multiview stereo can obtain accurate results. In the flat area, normal estimation networks usually obtain a good normal estimation. Based on the above observations, we reduce the possible spatial variation range of implicit neural surfaces by reliable geometric priors to alleviate shape-radiance ambiguity. Specifically, we use multiview stereo results to limit the NeuralRoom optimization space and then use reliable geometric priors to guide NeuralRoom training. Then the NeuralRoom would produce a neural scene representation that can render an image consistent with the input training images. In addition, we propose a smoothing method called perturbation-residual restrictions to improve the accuracy and completeness of the flat region, which assumes that the sampling points in a local surface should have the same normal and similar distance to the observation center. Experiments on the ScanNet dataset show that our method can reconstruct the texture-less area of indoor scenes while maintaining the accuracy of detail. We also apply NeuralRoom to more advanced multiview reconstruction algorithms and significantly improve their reconstruction quality.
Yusen Wang 0002, Zongcheng Li, Yu Jiang 0007, Kaixuan Zhou, Tuo Cao, Yanping Fu, Chunxia Xiao
ACM Trans. Graph.5
2021 Soudain: Online Adaptive Profile Configuration for Real-time Video Analytics
abstract
Since the real-time video analytics with high accuracy requirement is resource-consuming, the profiles regarding such resource-accuracy trade-off are needed before the analytics for better resource allocation at resource-constrained edges. With the inner changes of the video contents, outdated profiles fail to capture the trade-off dynamically over time, which requires the profiles to be updated periodically and incurs an overwhelming resource overhead. Thus, we present Soudain, which dynamically adjusts the configurations in profiles and corresponding profiling intervals to capture the inner changes of multiple video streams at edges. Upon the fine-grained decisions for profiles, we propose an integer program to maximize the accuracy of video analytics in a long-term scope with resource constraint, and then design an algorithm to adjust the profiles in an online manner. We implement Soudain upon the server with GPU. Our testbed evaluations confirm that, by using the live video streams derived from real-world traffic cameras, Soudain ensures the real-time requirement and achieves up to 25% improvement on the detection accuracy, compared with multiple state-of-the-art alternatives.
Yibo Jin 0001, Weiwei Miao, Zeng Zeng, Zhuzhong Qian, Jingmian Wang, Mingxian Zhou, Tuo Cao
IWQoS8
2021 Service Placement and Bandwidth Allocation for MEC-enabled Mobile Cloud Gaming
abstract
Mobile cloud gaming (MCG), which is potential to deliver high-quality gaming experience to users anywhere and anytime, suffers from tremendous wide-area traffic and long network delays. Mobile edge computing (MEC), where cloud computing capabilities are pushed to the network edge, can help by providing gaming services in the proximity to users. However, since the quality of experience (QoE), i.e., the gaming experience, is easily impaired by long network delays and low frame rates, the performance of MEC-enabled MCG highly depends on the placement of gaming services and the allocation of related bandwidth. Furthermore, due to the erratic mobility of users, migrating services to follow such mobility decreases the impairment but incurs extra system cost, leading to the performance-cost tradeoff. To address these challenges, in this paper, we jointly investigate service placement and bandwidth allocation for MEC-enabled MCG. Considering the system dynamics, we propose to minimize the QoE impairment in a long time scope under a cost constraint for long-term migrations. To solve the problem, we develop an online two-layer iterative algorithm OnTrial. Rigorous theoretical analyses demonstrate that OnTrial achieves a near-optimal performance and bounds the potential violation of the migration cost constraint. Simulation results show that OnTrial outperforms other algorithms by at least 25% on the long-term QoE impairment.
Tuo Cao, Zhuzhong Qian, Mingxian Zhou, Yibo Jin 0001
WOWMOM1