VLDB 2026 Research / reviewers in the wild / expert
Yutao Tang
dblp:128/0265
· DBLP profile ↗
20ranked-venue papers
9as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Sensor-Free and Explainable Method for Insulator Flashover Detection in Transmission Lines Based on Traveling Wave Data
Hongchun Shu, Yutao Tang, Jing Na, Xuan Su, Hongfang Zhao, Weizhong Sun, Weijie Lou, Yiming Han |
Adv. Eng. Informatics | 3 |
| 2025 | SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D ReconstructionabstractRecent efforts in Gaussian-Splat-based Novel View Synthesis can achieve photorealistic rendering; however, such capability is limited in sparse-view scenarios due to sparse initialization and over-fitting floaters. Recent progress in depth estimation and alignment can provide dense point cloud using few views; however, the resulting pose accuracy is suboptimal. In this work, we present SPARS3R, which combines the advantages of accurate pose estimation from Structure-from-Motion and dense point cloud from depth estimation. To this end, SPARS3R first performs a Global Fusion Alignment process that maps a prior dense point cloud to a sparse point cloud from Structure-from-Motion based on triangulated correspondences. RANSAC is applied during this process to distinguish inliers and outliers. SPARS3R then performs a second, Semantic Outlier Alignment step, which extracts semantically coherent regions around the outliers and performs local alignment in these regions. Along with several improvements in the evaluation process, we demonstrate that SPARS3R can achieve photorealistic rendering with sparse images and significantly outperforms existing approaches. Yutao Tang, Yuxiang Guo 0001, Deming Li, Cheng Peng 0008 |
CVPR | 1 |
| 2025 | RDFNet: Real-time Object Detection Framework for Foggy ScenesabstractDetecting objects in foggy scenes remains a persistent challenge, as detectors trained for fair weather often struggle with foggy data due to blurring effects. Previous methods based on domain adaptation, multi-task learning, etc. try to tackle this challenge, but they often fail to achieve an optimal balance between model complexity and accuracy. Hence, we propose a multi-branch pooling information fusion (MPIF) module, which combines local and global information to enhance feature representation with minimal computational overhead. We also design a lightweight multi-scale dehazing network (LMDNet) and utilize the multi-task learning strategy to adaptively incorporate dehazing feature information into the object detection network. Leveraging these core modules with additional design optimizations, we construct a novel real-time object detection framework, called RDFNet, for foggy images. Extensive experiments demonstrate that RDFNet outperforms SOTA detection methods for foggy scenes while enjoying less complexity and faster detection speeds. The source code will be released at https://github.com/PolarisFTL/RDFNet. Tianle Fang, Zhenbing Liu, Yutao Tang, Yingxin Huang, Haoxiang Lu, Chuangtao Zheng |
ICME | 3 |
| 2025 | MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the WildabstractIn-the-wild photo collections often contain limited volumes of imagery and exhibit multiple appearances, e.g., taken at different times of day or seasons, posing significant challenges to scene reconstruction and novel view synthesis. Although recent adaptations of Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have improved in these areas, they tend to oversmooth and are prone to overfitting. In this paper, we present MS-GS, a novel framework designed with \textbf{M}ulti-appearance capabilities in \textbf{S}parse-view scenarios using 3D\textbf{GS}. To address the lack of support due to sparse initializations, our approach is built on the geometric priors elicited from monocular depth estimations. The key lies in extracting and utilizing local semantic regions with a Structure-from-Motion (SfM) points anchored algorithm for reliable alignment and geometry cues. Then, to introduce multi-view constraints, we propose a series of geometry-guided supervision steps at virtual views in pixel and feature levels to encourage 3D consistency and reduce overfitting. We also introduce a dataset and an in-the-wild experiment setting to set up more realistic benchmarks. We demonstrate that MS-GS achieves photorealistic renderings under various challenging sparse-view and multi-appearance conditions, and outperforms existing approaches significantly across different datasets. Deming Li, Yutao Tang, Ravi Ramamoorthi, Rama Chellappa, Cheng Peng 0008 |
NeurIPS | 3 |
| 2024 | BAGS: Blur Agnostic Gaussian Splatting Through Multi-scale Kernel Modeling
Cheng Peng 0008, Yutao Tang, Nengyu Wang, Xijun Liu, Deming Li, Rama Chellappa |
ECCV (80) | 2 |
| 2024 | Sod-Uav: Small Object Detection For Unmanned Aerial Vehicle Images Via Improved Yolov7abstractDetecting small objects in Unmanned Aerial Vehicle (UAV) images is pivotal for a multitude of applications. Given the high-altitude perspective of UAVs, the images they capture often feature intricate backgrounds, pronounced object heterogeneity, and a plethora of sparsely situated small targets. These characteristics pose significant challenges to conventional detection algorithms. In response, we introduce an enhanced YOLOv7-based technique specifically tailored for small object detection in UAV images. Our approach adds more layers dedicated to small object detection and leverages the Bi-directional Feature Pyramid Network (BiFPN) to extract features across diverse scales. Furthermore, we enhance the detection heads by standardizing channel configurations and incorporate attention mechanisms through the dynamic head framework (DyHead). This allows the model to adaptively modify the detection head structure, catering to varying scales, tasks, and features. Preliminary results on the VisDrone2019 dataset indicate that our method surpasses existing state-of-the-art algorithms in UAV-based small object detection. Yifu Wang, Zihang Ma, Xinghe Wang, Yutao Tang |
ICASSP | 5 |
| 2024 | Semantic-aware Video Representation for Few-shot Action RecognitionabstractRecent work on action recognition leverages 3D features and textual information to achieve state-of-the-art performance. However, most of the current few-shot action recognition methods still rely on 2D frame-level representations, often require additional components to model temporal relations, and employ complex distance functions to achieve accurate alignment of these representations. In addition, existing methods struggle to effectively integrate textual semantics, some resorting to concatenation or addition of textual and visual features, and some using text merely as an additional supervision without truly achieving feature fusion and information transfer from different modalities. In this work, we propose a simple yet effective Semantic-Aware Few-Shot Action Recognition (SAFSAR) model to address these issues. We show that directly leveraging a 3D feature extractor combined with an effective feature-fusion scheme, and a simple cosine similarity for classification can yield better performance without the need of extra components for temporal modeling or complex distance functions. We introduce an innovative scheme to encode the textual semantics into the video representation which adaptively fuses features from text and video, and encourages the visual encoder to extract more semantically consistent features. In this scheme, SAFSAR achieves alignment and fusion in a compact way. Experiments on five challenging few-shot action recognition benchmarks under various settings demonstrate that the proposed SAFSAR model significantly improves the state-of-the-art performance. Yutao Tang, Benjamín Béjar Haro, René Vidal |
WACV | 1 |
| 2024 | An Improved Lightweight YOLO Algorithm for Recognition of GPS Interference Signals in Civil AviationabstractConsidering several sources that cause global position system (GPS) interference in civil aviation and the challenges faced by interference recognition algorithms in terms of efficiency and accuracy, we propose an improved You Only Look Once (YOLO)v7‐CHS algorithm (YOLOv7‐CHS) and investigate its effectiveness in identifying GPS signals and different types of interference signals. First, continuous wavelet transform (CWT) is introduced as a method for processing and analyzing signals in the time–frequency (TF) domain to effectively obtain their temporal and spectral characteristic information. Second, the ConvNeXt structure is integrated into the YOLOv7 backbone network to create a ConvNeXtBlock (CNeB) module to enhance the classification and recognition accuracy of interference signals. Additionally, an attention mechanism is introduced to further improve model recognition accuracy. To effectively improve the capability of signal feature extraction and mitigate the impact of background noise on TF feature suppression, we have integrated the efficient channel attention (ECA) channel attention module with the convolutional block attention module (CBAM) spatial attention module, thereby proposing a hybrid CBAM and ECA (HCE) attention module. Last, to address issues arising from accidental deletion of detection frames and multipath interference negatively affecting model recognition performance, we have employed the soft nonmaximum suppression (Soft‐NMS) algorithm while selecting an optimal loss function through comparative analysis. The comparative evaluation experimental results under different circumstances show that YOLOv7‐CHS achieves recognition accuracies of 98.0% and 99.6% for various types of signals, respectively. These values represent an increase of 1.7% and 1%, respectively, compared to YOLOv7. Moreover, in terms of lightweight indicators, YOLOv7‐CHS exhibits a significant improvement in performance: the frames per second (FPS) is increased by 75.1, the number of parameters (Params) was reduced by 4.75 M, and giga floating point operations per second (GFLOPs) were reduced by 65.9 G while effectively enhancing recognition capabilities. The proposed YOLOv7‐CHS not only improves signal recognition accuracy but also reduces model Params and computational complexity, achieving a lightweight model with promising application prospects in the rapid detection and recognition of GPS interference sources in civil aviation. Mian Zhong, Maonan Hu, Jiaqing Shen, Yutao Tang, Hede Lu |
IET Signal Process. | 6 |
| 2024 | Distributed Nash Equilibrium Seeking Dynamics With Discrete CommunicationabstractIn this brief, we aim to provide a distributed Nash equilibrium seeking algorithm in continuous time with discrete communications. A group of agents are considered playing a continuous-kernel noncooperative game over a network. The agents need to seek the Nash equilibrium when each player cannot get the overall action profiles in real time, rather are only able to get information from its networked neighbors. Meanwhile, a continuous-time dynamics is discussed for the players to update their variables, but the communications over the network are only assumed to allow at discrete-time instants, since continuous-time communications are prohibitive and cumbersome in practice. First, the periodic communication is considered at a fixed interval, and the solvability of Nash equilibrium seeking is shown with discrete communications. Then, an event-trigger communication scheme is proposed to further reduce the communication rounds. Nevertheless, the event-trigger communication scheme requires each player continuously monitoring its local states. To alleviate the monitoring burden, a periodic event detection mechanism is further developed. The exponential convergence of the dynamics with the three discrete communication schemes is proven. Finally, the comparative simulation studies are designed to illustrate the algorithm performance with different communication schemes and parameter settings. Rui Yu 0001, Yutao Tang, Peng Yi 0001, Li Li 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | A hierarchical growth method for extracting 3D phenotypic trait of apple tree branch in edge computing
Chengjian Zhang, Yutao Tang |
Wirel. Networks | 5 |
| 2022 | Facial Tic Detection in Untrimmed Videos of Tourette Syndrome PatientsabstractTourette Syndrome (TS) is a behavioral disorder that onsets in childhood and is characterized by the expression of involuntary movements and sounds commonly referred to as tics. Behavioral therapy is the first-line treatment for patients with TS, and it helps patients raise awareness about tic occurrence as well as develop tic inhibition strategies. However, the limited availability of therapists and the difficulties for in-home follow up work limits its effectiveness. An automatic tic detection system that is easy to deploy could alleviate the difficulties of home-therapy by providing feedback to the patients while exercising tic awareness. In this work, we propose a novel architecture (T-Net) for automatic tic detection and classification from untrimmed videos. T-Net combines temporal detection and segmentation and operates on features that are interpretable to a clinician. We compare T-Net to several state-of-the-art systems working on deep features extracted from the raw videos and T-Net achieves comparable performance in terms of average precision while relying on interpretable features needed in clinical practice. Yutao Tang, Benjamín Béjar Haro, Joey K.-Y. Essoe, Joseph F. McGuire, René Vidal |
ICPR | 1 |
| 2022 | vTrust: Remotely Executing Mobile Apps Transparently With Local Untrusted OSabstractIncreasingly, many security and privacy sensitive applications (apps for short) are running in the mobile platforms. However, as the mobile operating systems are becoming increasingly sophisticated, they are vulnerable to various attacks. In addressing the need of running high assurance mobile apps in a secure environment even though the operating systems are untrusted, this paper presents VTRUST, a new mobile app trusted execution environment, which offloads the general execution and storage of a mobile app to a trusted remote server (e.g., a VM running in a cloud) and secures the I/O between the server and the mobile device with the aid of a trusted hypervisor on the mobile device. Specifically, VTRUST establishes an encrypted I/O channel between the local hypervisor and the remote server, such that any sensitive data flowing through the mobile OS, which is hosted by the hypervisor, is encrypted from the perspective of the local mobile OS. To enhance the performance of VTRUST, we have also designed multiple optimizations, such as output data compression and selective sensor data transmission. We have implemented VTRUST and our evaluation shows that it has limited impact on both user experience and the app performance. Yutao Tang, Zhengrui Qin, Zhiqiang Lin 0001, Yue Li 0002, Shanhe Yi, Fengyuan Xu, Qun Li 0001 |
IEEE Trans. Computers | 1 |
| 2021 | User input enrichment via sensing devices
Yutao Tang, Yue Li 0002, Qun Li 0001, Kun Sun 0001, Haining Wang 0001, Zhengrui Qin |
Comput. Networks | 1 |
| 2020 | Distributed Optimal Steady-State Regulation for High-Order Multiagent Systems With External DisturbancesabstractIn this paper, a distributed optimal steady-state regulation problem is formulated and investigated for heterogeneous linear multiagent systems subject to external disturbances. We aim to steer this high-order multiagent network to a prescribed steady-state determined as the optimal solution of a resource allocation problem in a distributed way. To solve this problem, we employ an embedded control design and convert the formulated problem to two simpler subproblems. Then, both state-feedback and output feedback controls are presented under mild assumptions to solve this problem with disturbance rejection. Moreover, we extend these results to the case with only real-time gradient information by high-gain control techniques. Finally, the numerical simulations verify their effectiveness. Yutao Tang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2019 | Optimal Output Consensus of High-Order Multiagent Systems With Embedded TechniqueabstractIn this paper, we study an optimal output consensus problem for a multiagent network with agents in the form of multi-input multioutput minimum-phase dynamics. Optimal output consensus can be taken as an extended version of the existing output consensus problem for higher-order agents with an optimization requirement, where the output variables of agents are driven to achieve a consensus on the optimal solution of a global cost function. To solve this problem, we first construct an optimal signal generator, and then propose an embedded control scheme by embedding the generator in the feedback loop. We give two kinds of algorithms based on different available information along with both state feedback and output feedback, and prove that these algorithms with the embedded technique can guarantee the solvability of the problem for high-order multiagent systems under standard assumptions. Yutao Tang, Zhenhua Deng, Yiguang Hong |
IEEE Trans. Cybern. | 1 |
| 2018 | NodeMerge: Template Based Efficient Data Reduction For Big-Data Causality AnalysisabstractToday's enterprises are exposed to sophisticated attacks, such as Advanced Persistent Threats~(APT) attacks, which usually consist of stealthy multiple steps. To counter these attacks, enterprises often rely on causality analysis on the system activity data collected from a ubiquitous system monitoring to discover the initial penetration point, and from there identify previously unknown attack steps. However, one major challenge for causality analysis is that the ubiquitous system monitoring generates a colossal amount of data and hosting such a huge amount of data is prohibitively expensive. Thus, there is a strong demand for techniques that reduce the storage of data for causality analysis and yet preserve the quality of the causality analysis. To address this problem, in this paper, we propose NodeMerge, a template based data reduction system for online system event storage. Specifically, our approach can directly work on the stream of system dependency data and achieve data reduction on the read-only file events based on their access patterns. It can either reduce the storage cost or improve the performance of causality analysis under the same budget. Only with a reasonable amount of resource for online data reduction, it nearly completely preserves the accuracy for causality analysis. The reduced form of data can be used directly with little overhead. To evaluate our approach, we conducted a set of comprehensive evaluations, which show that for different categories of workloads, our system can reduce the storage capacity of raw system dependency data by as high as 75.7 times, and the storage capacity of the state-of-the-art approach by as high as 32.6 times. Furthermore, the results also demonstrate that our approach keeps all the causality analysis information and has a reasonably small overhead in memory and hard disk. Yutao Tang, Ding Li 0001, Zhichun Li, Mu Zhang 0001, Kangkook Jee, Xusheng Xiao, Zhenyu Wu 0003, Junghwan Rhee, Fengyuan Xu, Qun Li 0001 |
CCS | 1 |
| 2016 | MobiPlay: a remote execution based record-and-replay tool for mobile applicationsabstractThe record-and-replay approach for software testing is important and valuable for developers in designing mobile applications. However, the existing solutions for recording and replaying Android applications are far from perfect. When considering the richness of mobile phones' input capabilities including touch screen, sensors, GPS, etc., existing approaches either fall short of covering all these different input types, or require elevated privileges that are not easily attained and can be dangerous. In this paper, we present a novel system, called MobiPlay, which aims to improve record-and-replay testing. By collaborating between a mobile phone and a server, we are the first to capture all possible inputs by doing so at the application layer, instead of at the Android framework layer or the Linux kernel layer, which would be infeasible without a server. MobiPlay runs the to-be-tested application on the server under exactly the same environment as the mobile phone, and displays the GUI of the application in real time on a thin client application installed on the mobile phone. From the perspective of the mobile phone user, the application appears to be local. We have implemented our system and evaluated it with tens of popular mobile applications showing that MobiPlay is efficient, flexible, and comprehensive. It can record all input data, including all sensor data, all touchscreen gestures, and GPS. It is able to record and replay on both the mobile phone and the server. Furthermore, it is suitable for both white-box and black-box testing. Zhengrui Qin, Yutao Tang, Edmund Novak, Qun Li 0001 |
ICSE | 2 |
| 2015 | Physical media covert channels on smart mobile devicesabstractIn recent years mobile smart devices such as tablets and smartphones have exploded in popularity. We are now in a world of ubiquitous smart devices that people rely on daily and carry everywhere. This is a fundamental shift for computing in two ways. Firstly, users increasingly place unprecedented amounts of sensitive information on these devices, which paints a precarious picture. Secondly, these devices commonly carry many physical world interfaces. In this paper, we propose information leakage malware, specifically designed for mobile devices, which uses covert channels over physical "real-world" media, such as sound or light. This malware is stealthy; able to circumvent current, and even state-of-the-art defenses to enable attacks including privilege escalation, and information leakage. We go on to present a defense mechanism, which balances security with usability to stop these attacks. Edmund Novak, Yutao Tang, Zijiang Hao, Qun Li 0001, Yifan Zhang 0002 |
UbiComp | 2 |
| 2015 | SMOC: A secure mobile cloud computing platformabstractMobile devices are now ubiquitous in the modern world. In this paper, we propose a novel and practical mobile-cloud platform for smart mobile devices. Our platform allows users to run the entire mobile device operating system and arbitrary applications on a cloud-based virtual machine. It has two design fundamentals. First, applications can freely migrate between the user's mobile device and a backend cloud server. We design a file system extension to enable this feature, so users can freely choose to run their applications either in the cloud (for high security guarantees), or on their local mobile device (for better user experience). Second, in order to protect user data on the smart mobile device, we leverage hardware virtualization technology, which isolates the data from the local mobile device operating system. We have implemented a prototype of our platform using off-the-shelf hardware, and performed an extensive evaluation of it. We show that our platform is efficient, practical, and secure. Zijiang Hao, Yutao Tang, Yifan Zhang 0002, Edmund Novak, Nancy J. Carter, Qun Li 0001 |
INFOCOM | 2 |
| 2012 | Hierarchical control design of nonlinear systems based on approximate simulationabstractHierarchical control for a class of nonlinear systems is discussed using approximate simulation relation in this paper. An error bound is obtained under a modified approximate simulation function at first. The interface construction problem is solved for a class of nonlinear systems. Moreover, with a small-gain type condition, the abstraction result is extended to interconnected systems with a sufficient condition for approximate simulation when a part of the considered system can be removed by abstraction. Yutao Tang, Yiguang Hong |
ICARCV | 1 |