Jiashuo Yu

dblp:289/7338 · DBLP profile ↗
← Back
36ranked-venue papers
10as first author
36since 2021 · last 2026
0000-0002-3094-6687ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 19 · 4 first-author · 19 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Turbolearn: Harnessing Accurate and Line-Rate Deep Learning on Programmable Switches
Zhifan Jiang, Longlong Zhu, Jiashuo Yu, Linying Zheng, Chunming Wu 0001, Xiang Chen 0017
APNet3
2026 TurboLearn: Harnessing Accurate and Line-Rate Deep Learning on Programmable Switches
abstract
The intelligent data plane (IDP) embeds deep learning (DL) models on switches for line-rate traffic analysis, but hardware constraints often force simplified models, reducing accuracy, while complex models like Transformers remain undeployable. We present TurboLearn, which achieves high accuracy and line-rate performance by co-designing inference across the switch ASIC and switch OS on the same switch. TurboLearn uses three key techniques: (1) a hardware fast path for lightweight classification and a software normal path for complex models, (2) confidence-based selective inference that escalates only low-confidence packets, and (3) confidence-calibrated knowledge distillation, where the normal path teaches the fast path. On Intel Tofino2 switches across three real-world traffic tasks, TurboLearn supports models that existing IDPs cannot deploy, improves macro-F1 by up to 31.31%, and keeps over 90% of traffic on the fast path with zero throughput loss.
Zhifan Jiang, Longlong Zhu, Jiashuo Yu, Linying Zheng, Chunming Wu 0001, Xiang Chen 0017
APNet3
2026 APTMatch: Empowering Dynamic Rule Updating for Learning-based Packet Classification
Jiashuo Yu, Longlong Zhu, Zongye Lin, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001
ICC2
2026 HYDRA: A Hybrid Synthesizer for Asymmetric Mixture-of-Experts Communication Scheduling
Lida Liao, Qianxun Xu, Xuanwei Si, Hongyan Liu 0001, Jiashuo Yu, Zongye Lin, Qiaoling Hu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001
ICC6
2026 Proteus: Towards Accurate and Low-overhead In-Network Malicious Traffic Detection
abstract
Network intrusion detection systems (NIDS) are essential for web security by identifying and dropping malicious traffic. Existing in-network NIDS leverage the Tbps-level packet processing capability of programmable switches to achieve high-speed flow classification. They translate complex trained machine learning models to decision trees (DTs), where DTs are deployed on programmable switches via single-DT or multiple-DT deployment. However, they face a fundamental trade-off: single-DT deployment suffers from low classification accuracy due to over-pruning of trees, while multiple-DT deployment suffers from high overhead due to deploying multiple tree replicas. In this paper, we propose Proteus, an in-network malicious traffic detection system that achieves both high classification accuracy and low overhead. Its key idea is to split the original DT into critical and normal sub-trees, where these sub-trees have different impacts on overall accuracy. More precisely, Proteus first splits a DT into one critical and several normal sub-trees for adapting to the accuracy requirement and switch resource budgets. Second, it minimizes coordination overhead between sub-trees while ensuring full flow coverage via mixed-integer linear programming. Third, it dynamically reallocates or migrates sub-trees to adapt to changing resources by monitoring both classification accuracy and switch resource changes. Testbed experiments with 12.8 Tbps programmable switches show that Proteus improves classification accuracy, reduces switch resource consumption, and reduces classification latency.
Longlong Zhu, Linying Zheng, Qing Shu, Zedi Chen, Jiashuo Yu, Shaopeng Zhou, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001, Xiang Chen 0017
WWW5
2026 DeepConfig: A verifiable configuration generation framework for MAN overlays using LLMs
Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Jiashuo Yu, Lida Liao
Comput. Networks5
2026 VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
abstract
Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal evaluation system should provide insights to inform future developments of video generation. To this end, we present VBench++, a comprehensive benchmark suite that dissects "video generation quality" into specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. VBench++ has several appealing properties: 1) Comprehensive Dimensions: VBench++ comprises 16 dimensions in text-to-video generation (e.g., subject identity inconsistency, motion smoothness, temporal flickering, and spatial relationship, etc). The evaluation metrics with fine-grained levels reveal individual models' strengths and weaknesses. 2) Human Alignment: We also provide a dataset of human preference annotations to validate our benchmarks' alignment with human perception, for each evaluation dimension respectively. 3) Valuable Insights: We look into current models' ability across various evaluation dimensions, and various content types. We also investigate the gaps between video and image generation models. 4) Versatile Benchmarking: VBench++ is designed to evaluate a wide range of video generation tasks, including text-to-video and image-to-video. We introduce a high-quality Image Suite with an adaptive aspect ratio to enable fair evaluations across different image-to-video generation settings. Beyond assessing technical quality, VBench++ evaluates the trustworthiness of video generative models, providing a more holistic view of model performance. 5) Full Open-Sourcing: We fully open-source VBench++, including all prompts, the Image Suite, evaluation methods, generated videos, and human preference annotations.
Fan Zhang 0045, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma 0008, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang 0003, Yaohui Wang 0001, Ying-Cong Chen, Limin Wang 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 VEU-Bench: Towards Comprehensive Understanding of Video Editing
abstract
Widely shared videos on the internet are often edited. Recently, although Video Large Language Models (Vid-LLMs) have made great progress in general video understanding tasks, their capabilities in video editing understanding (VEU) tasks remain unexplored. To address this gap, in this paper, we introduce VEU-Bench (Video Editing Understanding Benchmark), a comprehensive benchmark that categorizes video editing components across various dimensions, from intra-frame features like shot size to inter-shot attributes such as cut types and transitions. Unlike previous video editing understanding benchmarks that focus mainly on editing element classification, VEU-Bench encompasses 19 fine-grained tasks across three stages: recognition, reasoning, and judging. To enhance the annotation of VEU automatically, we built an annotation pipeline integrated with an ontology-based knowledge base. Through extensive experiments with 11 state-of-the-art Vid-LLMs, our findings reveal that current Vid-LLMs face significant challenges in VEU tasks, with some performing worse than random choice. To alleviate this issue, we develop Oscars1, a VEU expert model fine-tuned on the curated VEU-Bench dataset. It outperforms existing open-source Vid-LLMs on VEU-Bench by over 28.3% in accuracy and achieves performance comparable to commercial models like GPT-4o. We also demonstrate that incorporating VEU data significantly enhances the performance of Vid-LLMs on general video understanding benchmarks, with an average improvement of 8.3% across nine reasoning tasks. The code and data are available at project page
Bozheng Li, Yongliang Wu, Jiashuo Yu, Licheng Tang, Jiawang Cao, Jay Wu
CVPR4
2025 DHC: Distributed Homomorphic Compression for Gradient Aggregation in Allreduce
abstract
Distributed training is critical for efficiently developing deep neural networks (DNNs) on tasks like image classification and natural language processing. However, as model and dataset sizes continue to grow, high communication overhead during gradient exchanges has become a major bottleneck in distributed training. Although existing homomorphic compression frameworks effectively reduce communication overhead, their reliance on centralized architectures makes them unsuitable for the mainstream decentralized AllReduce architecture. To address this, we propose DHC, a framework for homomorphic gradient compression in AllReduce architectures. Its key idea is HG-Sketch, which leverages multi-level index tables for direct in-network aggregation of compressed gradients, thereby eliminating additional computational overhead. Additionally, DHC introduces an index-sharing method to optimize memory usage on programmable switches. Furthermore, we establish an Integer Linear Programming (ILP) model to optimize the deployment strategy of programmable switches, further enhancing in-network aggregation capabilities. Experimental results demonstrate that DHC achieves a$3.8 \times$increase in aggregation speed and a$4.2 \times$improvement in aggregation throughput.
Lida Liao, Zhengli Lin, Longlong Zhu, Hongyan Liu 0001, Jiashuo Yu, Dong Zhang 0010, Chunming Wu 0001
ICC6
2025 P4Alex: A Scalable Range Matching Approach for Programmable Switches
abstract
Range matching (RM), a flexible primitive for implementing network applications on programmable switches, is often subject to limited TCAM resources. Consequently, existing RM approaches rely on SRAM/ALU-assisted data structures to extend TCAM capacity. However, they consume significant SRAM, ALU, and pipeline stages, which hinders the implementation of other primitives (e.g., basic forwarding). In this paper, we propose P4Alex, a scalable RM framework for programmable switches. The key idea is leveraging emerging learned index structures to support RM, which replaces the storage by model inference for lightweight and fixed index depth. Unfortunately, the learning index structure cannot be directly implemented on programmable switches due to hardware limitations (e.g., floating-point computation and no-loop operations). In response, we design several optimizations in P4Alex to make it deployable. We successfully implemented P4Alex on Intel Tofino switches. Experimental results show that, compared to existing RM approaches, P4Alex extends RM capabilities from tens of thousands to millions. At the same scales, P4Alex reduces TCAM and SRAM resource consumption by up to 89.9% and 43.7%, respectively, while increasing latency by only$0.48 \mu ~\mathrm{s}$.
Jiashuo Yu, Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001
ICC1
2025 VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
abstract
We present VRBench, the first long narrative video benchmark crafted for evaluating large models' multi-step reasoning capabilities, addressing limitations in existing evaluations that overlook temporal reasoning and procedural validity. It comprises 960 long videos (with an average duration of 1.6 hours), along with 8,243 human-labeled multi-step question-answering pairs and 25,106 reasoning steps with timestamps. These videos are curated via a multi-stage filtering process including expert inter-rater reviewing to prioritize plot coherence. We develop a human-AI collaborative framework that generates coherent reasoning chains, each requiring multiple temporally grounded steps, spanning seven types (e.g., event attribution, implicit inference). VRBench designs a multi-phase evaluation pipeline that assesses models at both the outcome and process levels. Apart from the MCQs for the final results, we propose a progress-level LLM-guided scoring metric to evaluate the quality of the reasoning chain from multiple dimensions comprehensively. Through extensive evaluations of 12 LLMs and 19 VLMs on VRBench, we undertake a thorough analysis and provide valuable insights that advance the field of multi-step reasoning.
Jiashuo Yu, Yue Wu 0013, Meng Chu, Zhifei Ren, Zizheng Huang, Pei Chu, Yinan He, Zhenxiang Li, Zhongying Tu, Conghui He, Yu Qiao 0001, Yali Wang 0001, Yi Wang 0074, Limin Wang 0002
ICCV1
2025 OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
abstract
Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains the capabilities of large language models during multimodal fine-tuning. However, the limited scale and diversity of current image-text interleaved data restrict the development of multimodal large language models. In this paper, we introduce OmniCorpus, a 10 billion-scale image-text interleaved dataset. Using an efficient data engine, we filter and extract large-scale high-quality documents, which contain 8.6 billion images and 1,696 billion text tokens. Compared to counterparts (e.g., MMC4, OBELICS), our dataset 1) has 15 times larger scales while maintaining good data quality; 2) features more diverse sources, including both English and non-English websites as well as video-centric websites; 3) is more flexible, easily degradable from an image-text interleaved format to pure text corpus and image-text pairs. Through comprehensive analysis and experiments, we validate the quality, usability, and effectiveness of the proposed dataset. We hope this could provide a solid data foundation for future multimodal model research.
Qingyun Li, Zhe Chen 0017, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, Guanzhou Chen 0004, Yinan He, Zhangwei Gao, Erfei Cui, Jiashuo Yu, Hao Tian 0006, Bin Wang 0065, Xingjian Wei, Wei Li 0320, Wenjian Zhang, Bo Zhang 0069, Pinlong Cai
ICLR11
2025 EffiMatch: Enabling Fast and Accurate Learning-based Packet Classification
abstract
Learning-based Packet Classification methods reduce memory overhead by using lightweight Recursive Model Index(RMI) structures to limit the search range, followed by linear matching. However, they face a trade-off: complex RMI structures achieve smaller search ranges but slow down lookup, while simpler ones are faster but require larger scans. In this paper, we propose EffiMatch, a parallel multi-model lookup architecture aimed at resolving the trade-off between RMI complexity and linear search range in learning-based index systems. We propose two key designs: 1) We design a partitioning strategy called Distribution-Distance Partitioning (DDP), which groups data points with similar trends into the same segment. Combined with parallel lookup, this reduces the linear search range while maintaining high lookup speed. 2) We propose a more fine-grained binarization method, Base-Index Representation (BI), which approximates floating-point operations using integers. This method further reduces the search range without increasing model complexity. Experimental results show that EffiMatch reduces the linear search range by 26.84% using lower-complexity RMI models, which improves lookup speed by up to 6× and reduces construction time by up to 4 orders of magnitude compared to state-of-the-art LPC methods.
Lida Liao, Jiashuo Yu, Longlong Zhu, Hongyan Liu 0001, Dong Zhang 0010, Xiang Chen 0017, Chunming Wu 0001
ICNP3
2025 Scaling Learning-based Packet Classification Hardware with NeuTree
Jiashuo Yu, Longlong Zhu, Linying Zheng, Dong Zhang 0010, Xiang Chen 0017
INFOCOM1
2025 Monica: Towards Scalable Distributed System Verification by Programmable Switch-Based Testing
abstract
Data correctness in distributed systems is ensured by data consistency, where consistency is achieved by consensus algorithms. To safeguard data consistency, current testing tools use stress testing methods to examine consensus algorithms. However, existing tools are unable to simulate the situation under high traffic and suffer from excessive verification time. In this paper, we propose Monica, a scalable and efficient verification framework. Its key idea is to leverage the programmable switch to verify consensus algorithms. Specifically, Monica provides a set of primitives that researchers can invoke. Then, the control server recognizes the primitives and automatically configures the data plane. After that, the programmable switch collaborates with the control server to complete the verification. Experimental results show that Monica can generate traffic at the rate of Tbps level while keeping the computational and memory consumption of the programmable switch under 11.87%. Compared to existing testing tools, Monica increases the verification speed by up to 3.13 times. Further, Monica improved accuracy by 35.71% in high-traffic scenarios over other tools.
Jiashuo Yu, Longlong Zhu, Dong Zhang 0010, Lida Liao, Rongbang Wu, Xiang Chen 0017, Chunming Wu 0001
IWQoS2
2025 Polyx: Accelerating Verification of Traffic Migration in Large-Scale BGP Networks
abstract
In BGP networks, traffic migration verification ensures the scalability and reliability of the network during configuration changes. However, previous approaches suffer from low scalability and high computational overhead. In this poster, we propose Polyx, a framework for accelerating verification of traffic migration in large-scale BGP networks. Its key idea is to leverage hardware parallelism with a deterministic serialization algorithm to enhance state machine techniques. We implement the Polyx prototype and evaluate it on our built testbed. The experimental results demonstrate that Polyx achieves up to 46× overall speedup, 36× in state machine construction, and 131× in equivalence verification with minimal FPGA resource usage.
Rongbang Wu, Longlong Zhu, Jiashuo Yu, Dong Zhang 0010, Hongyan Liu 0001, Zongye Lin, Lida Liao, Xiang Chen 0017, Chunming Wu 0001
IWQoS3
2025 Handling Data Plane Program Deployment Dynamics with High-Quality Generative Diffusion Models
abstract
Deploying data plane programs across the network is typically formulated as a mixed-integer programming task, leading to a long execution time. In response, existing studies carefully tailor heuristics for specific task properties such as objectives. However, they suffer from poor solution quality under dynamic task deployment since they overfit specific task properties. Recently, generative diffusion models have been widely adopted in network optimizations due to their strong adaptability and generalization. Accordingly, in this poster, we propose a diffusion model-based framework for data plane program deployment tasks. Our key idea is to leverage the reverse denoising process of diffusion models to react to dynamic task changes at runtime while maintaining high solution quality. Preliminary results on our testbed show that we reduce latency by 66.67% and resource overhead by 58.62% during dynamic deployment.
Longlong Zhu, Jiashuo Yu, Xiang Chen 0017, Qing Shu, Zedi Chen, Zhifan Jiang, Qun Huang 0001, Xuan Liu 0006, Dong Zhang 0010, Chunming Wu 0001
IWQoS2
2025 LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models
Yaohui Wang 0001, Xin Ma 0031, Shangchen Zhou, Yi Wang 0074, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang 0001, Yuwei Guo 0002, Tianxing Wu 0002, Chenyang Si, Yuming Jiang 0003, Cunjian Chen, Chen Change Loy, Bo Dai 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002
Int. J. Comput. Vis.9
2024 VBench: Comprehensive Benchmark Suite for Video Generative Models
abstract
Video generation has witnessed significant advance-ments, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal eval-uation system should provide insights to inform future de-velopments of video generation. To this end, we present VBench, a comprehensive benchmark suite that dissects “video generation quality” into specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. VBench has three appealing proper-ties: 1) Comprehensive Dimensions: VBench comprises 16 dimensions in video generation (e.g., subject identity in-consistency, motion smoothness, temporal flickering, and spatial relationship, etc.). The evaluation metrics with fine-grained levels reveal individual models' strengths and weaknesses. 2) Human Alignment: We also provide a dataset of human preference annotations to validate our benchmarks' alignment with human perception, for each evaluation dimension respectively. 3) Valuable Insights: We look into current models' ability across various evaluation dimensions, and various content types. We also investi-gate the gaps between video and image generation models. We will open-source VBench, including all prompts, evaluation methods, generated videos, and human preference an-notations, and also include more video generation models in VBench to drive forward the field of video generation.
Yinan He, Jiashuo Yu, Fan Zhang 0045, Chenyang Si, Yuming Jiang 0003, Yuanhan Zhang, Tianxing Wu 0002, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang 0001, Limin Wang 0002, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002
CVPR3
2024 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Yi Wang 0074, Kunchang Li 0002, Xinhao Li 0004, Jiashuo Yu, Yinan He, Guo Chen 0006, Baoqi Pei, Rongkun Zheng, Zun Wang 0001, Yansong Shi, Tianxiang Jiang, Jilan Xu, Hongjie Zhang 0002, Yifei Huang 0002, Yu Qiao 0001, Yali Wang 0001, Limin Wang 0002
ECCV (85)4
2024 SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction
abstract
Recently video generation has achieved substantial progress with realistic results. Nevertheless, existing AI-generated videos are usually very short clips ("shot-level'') depicting a single scene. To deliver a coherent long video ("story-level''), it is desirable to have creative transition and prediction effects across different clips. This paper presents a short-to-long video diffusion model, SEINE, that focuses on generative transition and prediction. The goal is to generate high-quality long videos with smooth and creative transitions between scenes and varying lengths of shot-level videos. Specifically, we propose a random-mask video diffusion model to automatically generate transitions based on textual descriptions. By providing the images of different scenes as inputs, combined with text-based control, our model generates transition videos that ensure coherence and visual quality. Furthermore, the model can be readily extended to various tasks such as image-to-video animation and autoregressive video prediction. To conduct a comprehensive evaluation of this new generative task, we propose three assessing criteria for smooth and creative transition: temporal consistency, semantic similarity, and video-text semantic alignment. Extensive experiments validate the effectiveness of our approach over existing methods for generative transition and prediction, enabling the creation of story-level long videos.
Yaohui Wang 0001, Lingjun Zhang, Shaobin Zhuang, Xin Ma 0031, Jiashuo Yu, Yali Wang 0001, Dahua Lin, Yu Qiao 0001, Ziwei Liu 0002
ICLR6
2024 InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
abstract
This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation. InternVid contains over 7 million videos lasting nearly 760K hours, yielding 234M video clips accompanied by detailed descriptions of total 4.1B words. Our core contribution is to develop a scalable approach to autonomously build a high-quality video-text dataset with large language models (LLM), thereby showcasing its efficacy in learning video-language representation at scale. Specifically, we utilize a multi-scale approach to generate video-related descriptions. Furthermore, we introduce ViCLIP, a video-text representation learning model based on ViT-L. Learned on InternVid via contrastive learning, this model demonstrates leading zero-shot action recognition and competitive video retrieval performance. Beyond basic video understanding tasks like recognition and retrieval, our dataset and model have broad applications. They are particularly beneficial for generating interleaved video-text data for learning a video-centric dialogue system, advancing video-to-text and text-to-video generation research. These proposed resources provide a tool for researchers and practitioners interested in multimodal video understanding and generation.
Yi Wang 0074, Yinan He, Yizhuo Li 0001, Kunchang Li 0002, Jiashuo Yu, Xin Ma 0031, Xinhao Li 0004, Guo Chen 0006, Yaohui Wang 0001, Ping Luo 0002, Ziwei Liu 0002, Yali Wang 0001, Limin Wang 0002, Yu Qiao 0001
ICLR5
2024 TupleRadar: Accelerating Tuple Space Search in Packet Classification by Learned Index
abstract
Tuple space search(TSS)-based packet classification is the keystone of network system. Previous studies accelerate TSS by partitioning tuples, combining trees and tuples, and merging tuples. However, they do not scale with the number of rules, resulting in a high memory footprint or update time. In this paper, we propose TupleRadar, a framework for accelerating TSS while ensuring low memory footprint and fast rule updates. Our key idea is to construct learned indexes for tuples, which inherently improve the lookup speed but ensure the advantages of TSS. Specifically, TupleRadar builds orderly hash table-based tuples and then constructs the updatable learned index. It provides a bounded memory footprint of the index structure as well. We have evaluated TupleRadar on multiple scales rule-sets. Experimental results show that TupleRadar outperforms previous solutions, reducing 46.66% lookup time and 61.53% memory footprint on average, by up to 86.70% and 88.95%. It also performs a competitive rule update speed.
Longlong Zhu, Jiashuo Yu, Kaiwei Huang, Zhengyan Zhou, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001
IWQoS2
2024 DOT: Towards Fast Decision Tree Packet Classification by Optimizing Rule Partitions
abstract
Packet classification is a crucial component of modern networks. Existing decision tree-based algorithms alleviate the rule replication problem caused by overlapping rules in the ruleset via rule partitioning. They partition the ruleset into multiple subsets based on rule characteristics to reduce rule overlaps. However, existing algorithms fail to address the overlap between rules in the same set, seriously decreasing speed and memory performance. In this paper, we propose DOT, a framework for optimizing rule partitions before constructing decision trees. Its key idea is to migrate rules in subsets based on rule overlaps and the features of heuristics used to construct trees, as well as reorganize rules aided by tuples. DOT finds out the migrated rule candidates using rule dependency graphs and heuristic features, then transforms the rule migration problem into an integer linear programming problem and solves for the optimal migration strategy. Further, we employ a tuple-assisted approach to accelerate rule matching. Experiments show that DOT enhances existing decision tree-based algorithms, improving lookup speed by 1.69 ×, reducing average 24.85% memory consumption and 31.03% decision tree depth.
Longlong Zhu, Jiashuo Yu, Linying Zheng, Dong Zhang 0010, Chunming Wu 0001
LCN3
2024 TransTuple: Toward Fast Packet Classification via Adaptive Tuple Replacement
abstract
Open vSwitch (OVS) is a widely used software switch in virtualized environments and software-defined networks. OVS uses tuple space search (TSS) for packet classification in the datapath, allowing fast network rule updates, but the increasing number of rules poses a classification performance challenge. To address this, existing methods incorporate decision trees with TSS to form a hybrid structure, enhancing classification speed. However, decision trees tend to overfit the initial ruleset, becoming unbalanced after rule updates and leading to a sharp decline in classification performance. In this paper, we propose TransTuple, a framework to optimize hybrid structures for fast packet classification under rule updates. The core idea of TransTuple is to identify bottleneck branches in decision trees that degrade performance and to replace them with lightweight tuples, providing better throughput under rule updates. These tuples maintain rules using hash tables, enabling fast updating and packet matching on bottleneck branches. We use TransTuple to optimize three state-of-the-art hybrid structured methods, i.e., CutTSS, TabTree, and MBitTree, achieving up to a 3.1x improvement in classification speed during rule updates.
Jiashuo Yu, Longlong Zhu, Rongbang Wu, Linying Zheng, Hongyan Liu 0001, Dong Zhang 0010, Chunming Wu 0001
SECON1
2024 Learning Music-Dance Representations Through Explicit-Implicit Rhythm Synchronization
abstract
Although audio-visual representation has been proven to be applicable in many downstream tasks, the representation of dancing videos, which is more specific and always accompanied by music with complex auditory contents, remains challenging and uninvestigated. Considering the intrinsic alignment between the cadent movement of the dancer and music rhythm, we introduceMuDaR, a novelMusic-DanceRepresentation learning framework to perform the synchronization of music and dance rhythms both in explicit and implicit ways. Specifically, we derive the dance rhythms based on visual appearance and motion cues inspired by the music rhythm analysis. Then the visual rhythms are temporally aligned with the music counterparts, which are extracted by the amplitude of sound intensity. Meanwhile, we exploit the implicit coherence of rhythms implied in audio and visual streams by contrastive learning. The model learns the joint embedding by predicting the temporal consistency between audio-visual pairs. The music-dance representation, together with the capability of detecting audio and visual rhythms, can further be applied to three downstream tasks: (a) dance classification, (b) music-dance retrieval, and (c) music-dance retargeting. Extensive experiments demonstrate that our proposed framework outperforms other self-supervised methods by a large margin.
Jiashuo Yu, Junfu Pu, Ying Cheng 0005, Rui Feng 0001, Ying Shan
IEEE Trans. Multim.1
2023 MINT: Empowering Multiple Flow Definition Query for Network-Wide Measurement
abstract
Network management tasks rely on precise and fine-grained network information to make correct and appropriate decisions. These tasks (e.g., DDoS detection) require network information with multiple flow definitions to better manage the network. However, the existing works mainly focus on the query of multiple flow definitions on a single switch, without a thoughtful solution for this query in network-wide measurement. In this paper, to address this problem, we overcome several challenges and propose MINT, a system that enables the query for multiple flow definitions in network-wide measurement. The key insights of MINT are: deploying MFSketch to measure multiple flow definitions information on the switch, cutting MFSketch into fixed-size slices, and using in-band telemetry (INT) to carry the slice to the analyzer. Therefore, after the analyzer collects and reorganizes the slices, network operators can query multiple flow definitions information of the whole network for various network management tasks. We implemented a prototype of MINT on a Barefoot Tofino switch. Experimental results show that MINT provides reliable transmission and consistency guarantees while only using switch resources comparable to state-of-the-art works, with less than 1% additional network overhead. Additionally, MFSketch provides accurate measurements for multiple flow definitions query, outperforming other solutions in both accuracy and F1 score.
Jiayi Cai, Zhengyan Zhou, Tingxin Sun, Jiashuo Yu, Longlong Zhu, Chengze Li, Dong Zhang 0010, Chunming Wu 0001
ICC4
2023 MiCuts: Combing Bit-Based Cutting and Splitting for Efficient Packet Classification
abstract
Packet classification is a crucial component in computer networking. To achieve high throughput and low memory consumption, existing solutions apply different heuristics in each construction stage to build efficient decision trees. However, previous studies divide the tree construction process based on the scale of rule subsets which is indirect to the performance goal, leading to massive rule replication and high tree depth. In this paper, we propose MiCuts, a fine-grained framework for packet classification with both high speed and low memory footprint. Its key idea is directly utilizing rule replication and tree depth to divide the tree-building process into three stages, each with suitable optimization goals. First, it partitions rules and builds shallow semi-trees without rule replication via selecting effective bits. Second, it transforms the switching problem of heuristics into an ILP problem and aims to minimize memory consumption while ensuring high lookup speed. Third, it merges some nodes to eliminate memory explosion caused by splitting, where MiCuts combines splitting and linear search. Extensive experimental results on ClassBench show that MiCuts outperforms state-of-the-art approaches, improving lookup speed by 1.71× while reducing memory footprint by 74.4% on average.
Longlong Zhu, Jiashuo Yu, Linying Zheng, Jinfeng Pan, Zhengyan Zhou, Hanze Chen, Dong Zhang 0010, Xiang Chen 0010, Chunming Wu 0001
ICC2
2023 Long-Term Rhythmic Video Soundtracker
abstract
We consider the problem of generating musical soundtracks in sync with rhythmic visual cues. Most existing works rely on pre-defined music representations, leading to the incompetence of generative flexibility and complexity. Other methods directly generating video-conditioned waveforms suffer from limited scenarios, short lengths, and unstable generation quality. To this end, we present Long-Term Rhythmic Video Soundtracker (LORIS), a novel framework to synthesize long-term conditional waveforms. Specifically, our framework consists of a latent conditional diffusion probabilistic model to perform waveform synthesis. Furthermore, a series of context-aware conditioning encoders are proposed to take temporal information into consideration for a long-term generation. Notably, we extend our model’s applicability from dances to multiple sports scenarios such as floor exercise and figure skating. To perform comprehensive evaluations, we establish a benchmark for rhythmic video soundtracks including the pre-processed dataset, improved evaluation metrics, and robust generative baselines. Extensive experiments show that our model generates long-term soundtracks with state-of-the-art musical quality and rhythmic correspondence. Codes are available at https://github.com/OpenGVLab/LORIS.
Jiashuo Yu, Yaohui Wang 0001, Xiao Sun 0001, Yu Qiao 0001
ICML1
2023 DTRadar: Accelerating Search Process of Decision Trees in Packet Classification
abstract
Packet classification is an essential part of computer networks. Existing algorithms propose a partition process to address the memory explosion problem of the decision tree algorithm caused by the huge number of rules with multiple fields. However, the search process requires traversing multiple trees generated by the partition, which reduces the search efficiency. The existing algorithms take simple approaches to optimize the search process, which is low efficiency or high hardware overhead. In this paper, we propose DTRadar, a framework for expediting the decision tree packet lookup process. Its key idea is building an abstract One-Big-Tree(OBT) for multiple decision trees by establishing the middle data structure. DTRadar considers each decision tree as a splittable tree and organizes these subtrees by intermediate data structures. Extensive experiments show that DTRadar benefits existing decision tree-based solutions in classification time by 61.60%, and the memory footprint only increased by 4.21% on average.
Jiashuo Yu, Longlong Zhu, Dong Zhang 0010, Chunming Wu 0001
ISCC1
2022 FROD: An Efficient Framework for Optimizing Decision Trees in Packet Classification
abstract
To perform efficient packet classification, decision tree-based methods conduct decision trees via hand-tuned heuristics. Then the performance testing and optimization are executed to ensure an excellent searching speed and space overhead. Specifically, when the performance is below expectation, existing solutions attempt to optimize the algorithms, such as conducting more sophisticated heuristics. However, reconstruction or adjustment for algorithms produces an intolerable time overhead due to the long optimization period, caused by uncertain performance benefits and high pre-processing time. In this paper, we propose FROD, an efficient framework for optimizing the decision trees directly in packet classification. FROD raises a meticulous evaluation to accurately appraise decision trees constructed by different heuristics. It then seeks out the bottleneck components via a lightweight heuristic. After that, FROD searches the optimal division for inferior components considering structural constraints and characteristics of traffic distribution. Evaluation on ClassBench shows that FROD benefits existing decision tree-based solutions in classification time by 41% and memory footprint by 19% on average, and reduces classification time by up to 64%.
Longlong Zhu, Jiashuo Yu, Jiayi Cai, Jinfeng Pan, Zhigao Li, Zhengyan Zhou, Dong Zhang 0010, Chunming Wu 0001
IWQoS2
2022 MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video Parsing
abstract
Recognizing and localizing events in videos is a fundamental task for video understanding. Since events may occur in auditory and visual modalities, multimodal detailed perception is essential for complete scene comprehension. Most previous works attempted to analyze videos from a holistic perspective. However, they do not consider semantic information at multiple scales, which makes the model difficult to localize events in different lengths. In this paper, we present a Multimodal Pyramid Attentional Network (MM-Pyramid ) for event localization. Specifically, we first propose the attentive feature pyramid module. This module captures temporal pyramid features via several stacking pyramid units, each of them is composed of a fixed-size attention block and dilated convolution block. We also design an adaptive semantic fusion module, which leverages a unit-level attention block and a selective fusion block to integrate pyramid features interactively. Extensive experiments on audio-visual event localization and weakly-supervised audio-visual video parsing tasks verify the effectiveness of our approach.
Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang
ACM Multimedia1
2022 Modality-aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence Detection
abstract
Weakly-supervised audio-visual violence detection aims to distinguish snippets containing multimodal violence events with video-level labels. Many prior works perform audio-visual integration and interaction in an early or intermediate manner, yet overlooking the modality heterogeneousness over the weakly-supervised setting. In this paper, we analyze the modality asynchrony and undifferentiated instances phenomena of the multiple instance learning (MIL) procedure, and further investigate its negative impact on weakly-supervised audio-visual learning. To address these issues, we propose a modality-aware contrastive instance learning with self-distillation (MACIL-SD) strategy . Specifically, we leverage a lightweight two-stream network to generate audio and visual bags, in which unimodal background, violent, and normal instances are clustered into semi-bags in an unsupervised way. Then audio and visual violent semi-bag representations are assembled as positive pairs, and violent semi-bags are combined with background and normal instances in the opposite modality as contrastive negative pairs. Furthermore, a self-distillation module is applied to transfer unimodal visual knowledge to the audio-visual model, which alleviates noises and closes the semantic gap between unimodal and multimodal features. Experiments show that our framework outperforms previous methods with lower complexity on the large-scale XD-Violence dataset. Results also demonstrate that our proposed approach can be used as plug-in modules to enhance other networks. Codes are available at https://github.com/JustinYuu/MACIL_SD.
Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang
ACM Multimedia1
2021 Improving Multimodal Speech Enhancement by Incorporating Self-Supervised and Curriculum Learning
abstract
Speech enhancement in realistic scenarios still remains many challenges, such as complex background signals and data limitations. In this paper, we present a co-attention based framework that incorporates self-supervised and curriculum learning to derive the target speech in noisy environments. Specifically, we first leverage self-supervision to pre-train the co-attention model on the task of audio-visual synchronization. The pre-trained model can focus on the lip of speakers automatically, and then the self-supervised features from the model are combined with a u-net regression network to separate the spectrograms of sound mixtures. To make the training process easier and further improve the performance, we introduce the curriculum learning scheme for the training stage of speech enhancement. Extensive experiments show that our model achieves superior performance over previous self-supervised method for speech enhancement, and demonstrate the generalizability of our approach to the transferred dataset.
Ying Cheng 0005, Mengyu He, Jiashuo Yu, Rui Feng 0001
ICASSP3
2021 MPN: Multimodal Parallel Network for Audio-Visual Event Localization
abstract
Audio-visual event localization aims to localize an event that is both audible and visible in the wild, which is a widespread audio-visual scene analysis task for unconstrained videos. To address this task, we propose a Multimodal Parallel Network (MPN), which can perceive global semantics and unmixed local information parallelly. Specifically, our MPN framework consists of a classification subnetwork to predict event categories and a localization subnetwork to predict event boundaries. The classification subnetwork is constructed by the Multimodal Co-attention Module (MCM) and obtains global contexts. The localization subnetwork consists of Multimodal Bottleneck Attention Module (MBAM), which is designed to extract fine-grained segment-level contents. Extensive experiments demonstrate that our framework achieves the state-of-the-art performance both in fully supervised and weakly supervised settings on the Audio-Visual Event (AVE) dataset.
Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001
ICME1
2021 Exploring Logical Reasoning for Referring Expression Comprehension
abstract
Referring expression comprehension aims to localize the target object in an image referred by a natural language expression. Most existing approaches neglect the implicit logical correlations among fine-grained cues, e.g., categories, attributes, which are beneficial for distinguishing objects. In this paper, we propose a logic-guided approach to explore logical knowledge for referring expression comprehension in a hierarchical modular-based framework. Specifically, we propose to extract fine-grained cues in visual and textual domains and perform logical reasoning over them with explicit logical expressions to regularize the matching process without extra parameters. Besides, we propose to improve existing modular-based methods by introducing context information of objects in the relationship module. Extensive experiments are conducted on three referring expression datasets, and the results demonstrate that our model can produce more consistent predictions and further achieve superior performance compared with previous methods.
Ying Cheng 0005, Jiashuo Yu, Yuejie Zhang, Rui Feng 0001
ACM Multimedia3