VLDB 2026 Research / reviewers in the wild / expert
Weitao Tang
dblp:241/6063
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Orion: Steering Personalized Web Agents via Global-Micro Profiling and Adaptive Intent TrackingabstractRecently, Large Language Models (LLMs) based Web Agents have shown significant potential in web understanding and interaction tasks. However, their personalization ability and user experience remain limited by the ambiguity and dynamic nature of user intent, struggling to model diverse user interests and track intent changes over time. To address these challenges, this paper proposes Orion, a novel personalized Web Agent. Orion adopts a global-micro profiling mechanism to balance users' long-term stable preferences and scenario-based needs, and introduces context-aware interest retrieval to enhance personalization. Additionally, we design adaptive profile tracking and proactive disambiguation mechanisms to effectively address the continuous evolution of user intent in multi-turn interactions. Orion is optimized through end-to-end online reinforcement learning, improving personalized reasoning and decision-making ability in real interactive scenarios. Experiments demonstrate that Orion significantly outperforms state-of-the-art baselines in personalized understanding and task efficiency. Die Hu 0004, Jingguo Ge, Weitao Tang, He Kong 0003, Liangxiong Li, Bingzhen Wu |
AAAI | 3 |
| 2026 | Blazer: Encrypted Video Traffic Identification for Mixed Segment Transmission Pattern based on LLMabstractDetermining the source of encrypted video traffic is an important task in network regulation. In the context of Dynamic Adaptive Streaming over HTTP (DASH), the newly emerged mixed segment transmission pattern introduces substantial difficulties for fingerprint matching, especially under adverse network conditions. To address these challenges, we propose Blazer, a DASH encrypted video traffic identification method for the mixed segment transmission pattern. First, we design a novel fingerprint that integrates video and audio segment sequences. Then, we extract the traffic fingerprint from the TLS record layer of video traffic. Finally, by observing implicit segment-mixing constraints, we design a targeted prompt and Retrieval Augmented Generation (RAG) that enables Large Language Models (LLMs) to perform fingerprint matching effectively. Across 12 network scenarios, Blazer delivers substantially better performance than the other 4 SOTA methods. Weitao Tang, Meijie Du, Die Hu 0004, Zhao Li 0010, Rong Yang 0008, Qingyun Liu 0001 |
ICMR | 1 |
| 2025 | WebSurfer: Enhancing LLM Agents with Web-Wise Feedback for Web NavigationabstractAs the Internet’s complexity and information volume surge, the need for efficient web automation becomes critical. Traditional web agents struggle with redundant web content, which disrupts their understanding of the environment. They also face inefficiencies in multi-task scenarios due to handcrafted exemplars and encounter error accumulation in long-horizon tasks, exacerbated by web-specific complexities like nested structures and interactive elements. To address these issues, we introduce WebSurfer, a novel web agent designed to filter, learn, and adapt in complex environments. WebSurfer refines task-oriented states for clearer observations and employs an exemplar retrieval and ordering strategy to enhance LLMs’ understanding and adaptability to current tasks. Notably,WebSurfer features a novel web-wise insight feedback mechanism that enables continuous adaptation and strategy refinement. Evaluations demonstrate that WebSurfer outperforms state-of-the-art (SOTA) methods on realistic tasks, achieving higher accuracy and enhancing longterm adaptability. Die Hu 0004, Jingguo Ge, Weitao Tang, Guoyi Li, Liangxiong Li, Bingzhen Wu |
ICASSP | 3 |
| 2025 | Pioneer: Encrypted Video Traffic Identification for Mixed Transmission of Video-Audio SegmentsabstractThe spread of harmful content via video has made video traffic identification crucial for network regulation. In the new transmission mode, audio and video segments are mixed to combine into video chunks. However, in poor networks, such combination is unstable, and video chunks may be lost and retransmitted. To address these challenges, this paper proposes Pioneer, an encrypted video traffic identification method for mixed transmission of audio and video segments. We introduce a precise video chunk reconstruction method for video traffic encrypted by both TLS and QUIC. Additionally, we propose Pseudo-Siamese Attention-Convolutional Network (PSACN) to calculate the similarity between traffic and video, leveraging contrastive learning during training to mitigate the impact of poor networks. Pioneer significantly improves accuracy compared with state-of-the-art (SOTA) methods under various network environments. Notably, this is the first study to address this emerging new transmission mode. Weitao Tang, Taizhong Xu, Meijie Du, Die Hu 0004, Qingyun Liu 0001 |
ICME | 1 |
| 2025 | Contrastive graph auto-encoder for graph embedding
Shuaishuai Zu, Li Li 0006, Jun Shen 0001, Weitao Tang |
Neural Networks | 4 |
| 2024 | GuessKT: Improving Knowledge Tracing via Considering Guess BehaviorsabstractKnowledge tracing (KT) aims to predict students’ responses to given questions based on their historical question-answering interactions. Recent studies have proposed multiple types of KT models, mainly relying on learners’ feedback to capture the evolution of their knowledge states. However, these models leave the influence of learners’ guess behaviors out of consideration. These behaviors would cause unreliable feedback and mislead the inference process of the KT models. Exploring the guess behaviors is important since it can potentially help us understand realistic learning interactions for more accurate knowledge state estimations. In this paper, we propose GuessKT to better trace the learning progress by introducing the guess behaviors into learning feedback attribution, leading to improved prediction performance. Specifically, a windowed attention network is proposed to capture rich information from local interactions, which enhances the reliability of extracted historical feedback. Furthermore, a recovery network is proposed to recover the students’ responses after additional mask processing, which bolsters the model’s ability to recognize guess behaviors. Experiments on five datasets show that our model GuessKT advanced predictive performance over other baselines. Shuaishuai Zu, Songtao Cai, Weitao Tang, Li Li 0006, Jun Shen 0001 |
ICASSP | 3 |
| 2024 | TSIV: A Two-Stage Approach for Identifying Encrypted Video Traffic in Unstable Network
Die Hu 0004, Jingguo Ge, Tong Li 0012, Hui Li 0098, Liangxiong Li, Weitao Tang |
ICONIP (6) | 6 |
| 2024 | Zenith: Real-time Identification of DASH Encrypted Video Traffic with DistortionabstractSome video traffic carries harmful content, such as hate speech and child abuse, primarily encrypted and transmitted through Dynamic Adaptive Streaming over HTTP (DASH). Promptly identifying and intercepting traffic of harmful videos is crucial in network regulation. However, QUIC is becoming another DASH transport protocol in addition to TCP. On the other hand, complex network environments and diverse playback modes lead to significant distortions in traffic. The issues above have not been effectively addressed. This paper proposes a real-time identification method for DASH encrypted video traffic with distortion, named Zenith. We extract stable video segment sequences under various itags as video fingerprints to tackle resolution changes and propose a method of traffic fingerprint extraction under QUIC and VPN. Subsequently, simulating the sequence matching problem as a natural language problem, we propose Traffic Language Model (TLM), which can effectively address video data loss and retransmission. Finally, we propose a frequency dictionary to accelerate Zenith's speed further. Zenith significantly improves accuracy and speed compared to other SOTA methods in various complex scenarios, especially in QUIC, VPN, automatic resolution, and low bandwidth. Zenith requires traffic for just half a minute of video content to achieve precise identification, demonstrating its real-time effectiveness. Weitao Tang, Meijie Du, Die Hu 0004, Qingyun Liu 0001 |
ACM Multimedia | 1 |
| 2023 | Long-Short Terms Frequency: A Method for Encrypted Video Streaming IdentificationabstractNowadays, with the vigorous development of self-media services, more and more individual users upload videos freely. While bringing goodness, it also inevitably brings evil. Therefore, it is particularly necessary to identify and supervise illegal videos through network stream. However, many video streaming services, such as YouTube, have applied encryption to protect users’ privacy, which makes it more difficult to analyze network stream. Many researches show that DASH (Dynamic Adaptive Streaming over HTTP) will leak information about video segmentation, which is related to the video content. Consequently, it is possible to analyze the content of encrypted video stream without decryption. Previous studies have proposed a series of encrypted video identification methods based on this. However, most of them need to wait for a long video playback time, such as more than 10s, or even wait for the entire video playback to complete the identification. In this paper, we propose a fast, lightweight, and accurate method named Long-Short Terms Frequency(LSTF) for online encrypted video identification. Experiments have proved that compared with the state-of-the-art, our method has advantages in both speed and accuracy, and even if CDN switching occurs during the video playback, it still has a high identification accuracy. Meijie Du, Minchao Xu, Kedong Liu, Weitao Tang, Lijuan Zheng, Qingyun Liu 0001 |
CSCWD | 4 |
| 2023 | Shrink: Identification of Encrypted Video Traffic Based on QUICabstractWith the increasing prevalence of network videos, video traffic has become a significant portion of overall network traffic. Due to the presence of harmful content such as pornography and violence in network videos, network monitoring is necessary. However, the encryption of videos poses challenges for network monitoring. More and more video service providers are adopting QUIC as the default video transmission protocol to accelerate data transfer speeds. However, the existing methods for identifying encrypted video traffic do not apply to QUIC. Video service providers typically employ Content Delivery Network (CDN) technology to enhance user experience, which can result in missing video chunks for side-channel identification. Additionally, fluctuations in network conditions can lead to the retransmission of video chunks. This paper proposes Shrink, a QUIC-based encrypted video traffic identification method. It effectively extracts video chunks from online QUIC encrypted video traffic and proposes a bucket structure and global-local match to alleviate the issues of video chunks retransmission and loss. Furthermore, a bucket word dictionary is designed to enhance the method’s running speed. Experimental results demonstrate that Shrink performs well in real network environments, exhibiting superior accuracy and speed compared to existing state-of-the-art methods. Weitao Tang, Meijie Du, Zhao Li 0010, Zhou Zhou 0007, Qingyun Liu 0001 |
IPCCC | 1 |
| 2021 | Anomaly detection of core failures in die casting X-ray inspection images using a convolutional autoencoder
Weitao Tang, Corey M. Vian, Ziyang Tang, Baijian Yang 0001 |
Mach. Vis. Appl. | 1 |
| 2020 | Low-Rank Sparse Tensor Approximations for Large High-Resolution VideosabstractTensor decomposition techniques are becoming increasingly important in processing videos with large sizes and dimensions. Under the framework of CANDECOMP/PARAFAC decomposition (CPD), this work studies low-rank sparse tensor approximations (LRSTAs) to higher-order tensors. Both theoretical and practical properties are evaluated for LRSTAs to represent large high-resolution videos. The evaluation brings three major contributions of this work. Firstly, the theoretical connection between CPD for high-order tensors and traditional singular value decomposition (SVD) for matrices are established, and the tensor rank for traditional SVD is defined. This provides a theoretical basis to compare tensor-based approach against matrix-based approach under the framework of tensor decompositions. Secondly, the non-orthogonality of CPD and its implications are revealed. The solution set of an LRSTA can only be used as a whole. Thirdly, a computationally efficient algorithm is developed. Its practical properties are also investigated in object detection and recognition in high-resolution videos. The results of the experiments showed that the proposed algorithm can handle large high-resolution videos very efficiently in terms of memory allocation. Results also revealed that commonly used total variations may not be a good evaluation metric for real world applications in computer vision. LRSTAs should be evaluated using the end goal of the applications, such as the accuracy of object detection and recognition. Xiang Liu 0016, Huyunting Huang, Weitao Tang, Tonglin Zhang, Baijian Yang 0001 |
ICMLA | 3 |