Shaoteng Liu

dblp:02/10511 · DBLP profile ↗
← Back
38ranked-venue papers
15as first author
25since 2021 · last 2026
0000-0001-5407-2905ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 8 since 2021Systems, architecture and hardware · 9 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 5 since 2021Computer networks · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
Jinwu Yang, Jiaan Wu, Xinyang Ma, Hairui Zhao 0002, Yida Gu, Yuanhong Huang, Wenjing Huang 0002, Yili Ma, Zhongzhe Hu, Shaoteng Liu, Jiaxun Lu, Guangming Tan, Dingwen Tao
ISCA16
2026 Generic Construction of Optimal-Access Binary MDS Array Codes with Smaller Sub-packetization
abstract
A $(k+r,k,l)$ binary array code of length $k+r$, dimension $k$, and sub-packetization $l$ is composed of $l\times(k+r)$ matrices over $\mathbb{F}_2$, with every column of the matrix stored on a separate node in the distributed storage system and viewed as a coordinate of the codeword. It is said to be maximum distance separable (MDS) if any $k$ out of $k+r$ coordinates suffice to reconstruct the whole codeword. The repair problem of binary MDS array codes has drawn much attention, particularly for single-node failures. In this paper, given an arbitrary binary MDS array code with sub-packetization $m$ as the base code, we propose two generic approaches (Generic Construction I and II) for constructing binary MDS array codes with optimal access (or repair) bandwidth for single-node failures. For every $s\leq r$, a $(k+r,k,ms^{\lceil \frac{k+r}{s}\rceil})$ code $\mathcal{C}_1$ with optimal access bandwidth can be constructed by Generic Construction I. Repairing a failed node of $\mathcal{C}_1$ requires connecting to $d = k+s-1$ helper nodes, in which $s-1$ helper nodes are designated and $k$ are free to select. $\mathcal{C}_1$ generally achieves smaller sub-packetization and provides greater flexibility in the selection of its coefficient matrices. For even $r\geq4$ and $s=\frac{r}{2}$ such that $s+1$ divides $k+r$, a $(k+r, k,ms^{\frac{k+r}{s+1}})$ code $\mathcal{C}_2$ with optimal repair bandwidth can be constructed by Generic Construction II, with $\frac{s}{s+1}(k+r)$ out of $k+r$ nodes having the optimal access property. To the best of our knowledge, $\mathcal{C}_2$ possesses the smallest sub-packetization among existing binary MDS array codes with optimal repair bandwidth known to date.
Qifu Tyler Sun, Shaoteng Liu, Liyang Zhou
ISIT3
2026 Access-friendly MDS Array Codes with Small Sub-packetization and Multiple Repair Degrees
Qifu Tyler Sun, Shaoteng Liu, Liyang Zhou
ISIT3
2026 Instantly Decodable Network Coding with Limited Feedback
Limin Wen, Rina Su, Qifu Tyler Sun, Shaoteng Liu
ISIT4
2026 Balanced Sparse Tree: A Scalable Network Topology for Large Language Models
abstract
The development of large language models (LLMs) has catalyzed unprecedented demand on the computing network, specifically for large-scale, few-hops, and low-latency, which directly underpin LLM task efficiency. However, mainstream topologies such as Clos suffer from costs and latency, while topologies with good scalability have symmetric or collective communication issues. In order to achieve a favorable balance among design metrics, we propose a novel topology named the Balanced Sparse Tree (BST), which is a topology characterized by symmetric design and sparse connections, motivated by hypergraph theory and Steiner Systems. Its degree-diameter upper-bound approaches the Moore Bound for Bipartite Biregular graphs, larger than other known dia-meter-2 topologies. Furthermore, we incorporate differentiated routing, deadlock freedom, and topology-affined deployment into BST. Testbed experiments, simulations, together with modeling analysis, demonstrate the superiority of BST over the state-of-the-art in network scale, latency, bandwidth, and cost. With equivalent scales, BST outperforms Clos with a 50% cost reduction while maintaining comparable performance for AI workloads. Furthermore, BST delivers a 3.9%–11.8% gain in collective communications and has 13.4% improvement over state-of-the-art topologies.
Shaoteng Liu, Dejun Kong 0001, Huitian Wang, Hongji Dong, Fuguang Huang, Xiaofeng Gao 0001, Bingyang Liu, Guihai Chen
SIGCOMM1
2026 Mini-Gemini: Mining the Potential of Multi-Modality Vision Language Models
abstract
In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic visual dialog and reasoning, a performance gap persists compared to advanced models like GPT-4 and Gemini. We propose a novel approach to narrow the gap by mining the potential of VLMs for better performance across various cross-modal tasks. It tackles the following questions: (1) How can high-resolution visual tokens improve image understanding without lengthening the token sequence? (2) How to improve reasoning and generation abilities of VLM with high-quality data? (3) How to close the gap between open-source VLMs and proprietary models on reasoning-driven generation? In particular, to enhance visual tokens, we propose to utilize an additional visual encoder for high-resolution refinement without increasing the visual token count. We further construct a high-quality dataset that promotes precise image comprehension and reasoning-based generation, expanding the operational scope of current VLMs. In general, Mini-Gemini further mines the potential of VLMs and empowers current frameworks with image understanding, reasoning, and generation simultaneously. The proposed model supports a series of dense and MoE Large Language Models (LLMs) from 2B to 34B, which achieve leading performance in several zero-shot benchmarks and even surpasses the developed private models. It is demonstrated to attain 80.6% accuracy on the MMB benchmark (+5.4 vs Gemini Pro) and 74.1% on TextVQA (+4.6 vs LLaVA-NeXT), achieving leading performance in several zero-shot benchmarks and even surpasses the developed private models. Furthermore, Mini-Gemini is proven to improve consistently with stronger LLM, visual encoder, and data in experiments.
Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Ruihang Chu, Shaoteng Liu, Jiaya Jia
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 Generative Video Propagation
abstract
Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of such models. Specifically, our framework, Gen-Prop, encodes the original video with a selective content encoder and propagates the changes made to the first frame using an image-to-video generation model. We propose a data generation scheme to cover multiple video tasks based on instance-level video segmentation datasets. Our model is trained by incorporating a mask prediction decoder head and optimizing a region-aware loss to aid the encoder to preserve the original content while the generation model propagates the modified region. This novel design opens up new possibilities: In editing scenarios, GenProp allows substantial changes to an object’s shape; for insertion, the inserted objects can exhibit independent motion; for removal, GenProp effectively removes effects like shadows and reflections from the whole video; for tracking, GenProp is capable of tracking objects and their associated effects together. Experiment results demonstrate the leading performance of our model in various video tasks, and we further provide in-depth analyses of the proposed framework.
Shaoteng Liu, Tianyu Wang 0003, Jui-Hsien Wang, Qing Liu 0017, Joon-Young Lee, Yijun Li 0001, Bei Yu 0001, Zhe Lin 0001, Soo Ye Kim, Jiaya Jia
CVPR1
2025 MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
abstract
This work explores enabling Chain-of-Thought (CoT) reasoning to link visual cues across multiple images. A straightforward solution is to adapt rule-based reinforcement learning for Vision-Language Models (VLMs). However, such methods typically rely on manually curated question-answer pairs, which can be particularly challenging when dealing with fine-grained visual details and complex logic across images. Inspired by self-supervised visual representation learning, we observe that images contain inherent constraints that can serve as supervision. Based on this insight, we construct image triplets comprising two augmented views of the same image and a third, similar but distinct image. During training, the model is prompted to generate a reasoning process to compare these images (i.e., determine same or different). Then we optimize the model with rule-based reinforcement learning. Due to the high visual similarity and the presence of augmentations, the model must attend to subtle visual cues and perform logical reasoning to succeed. Experimental results demonstrate that, although trained solely on visual comparison tasks, the learned reasoning ability generalizes effectively to a wide range of questions. Without relying on any human-annotated question-answer pairs, our method achieves significant improvements on multi-image reasoning benchmarks and shows strong performance on general vision tasks.
Xi Chen 0119, Mingkang Zhu, Shaoteng Liu, Xiaoyang Wu 0002, Xiaogang Xu 0002, Yu Liu 0063, Xiang Bai, Hengshuang Zhao
NeurIPS3
2025 Training-Free Efficient Video Generation via Dynamic Token Carving
abstract
Despite the remarkable generation quality of video Diffusion Transformer (DiT) models, their practical deployment is severely hindered by extensive computational requirements. This inefficiency stems from two key challenges: the quadratic complexity of self-attention with respect to token length and the multi-step nature of diffusion models. To address these limitations, we present Jenga, a novel inference pipeline that combines dynamic attention carving with progressive resolution generation. Our approach leverages two key insights: (1) early denoising steps do not require high-resolution latents, and (2) later steps do not require dense attention. Jenga introduces a block-wise attention mechanism that dynamically selects relevant token interactions using 3D space-filling curves, alongside a progressive resolution strategy that gradually increases latent resolution during generation. Experimental results demonstrate that Jenga achieves substantial speedups across multiple state-of-the-art video diffusion models while maintaining comparable generation quality (8.83$\times$ speedup with 0.01\% performance drop on VBench). As a plug-and-play solution, Jenga enables practical, high-quality video generation on modern hardware by reducing inference time from minutes to seconds---without requiring model retraining.
Yuechen Zhang, Jinbo Xing, Bin Xia 0014, Shaoteng Liu, Bohao Peng, Xin Tao 0001, Pengfei Wan 0001, Eric Lo 0001, Jiaya Jia
NeurIPS4
2025 Lightweight Instantly Decodable Network Coding: Performance Analysis and Algorithm Design
abstract
We consider broadcasting a block of data packets to multiple users via instantly decodable network coding (IDNC) under the semi-online feedback transmission mode. In this paper, we first introduce a new class of IDNC schemes called lightweight IDNC, tailored for wireless broadcast with stringent computational load at the receiver end. Unlike traditional IDNC that may encode a larger number of original packets together, lightweight IDNC limits each coded packet to a combination of at most two original packets. Explicit lower bounds of the total completion delay as well as the decoding delay are respectively obtained for arbitrary lightweight IDNC schemes. We further investigate the number of transmission rounds as another performance metric, and explicitly characterize its distribution and expectation. The characterizations apply to arbitrary partition-based IDNC schemes, including the lightweight IDNC schemes considered in this paper. A new efficient algorithm is also proposed to construct lightweight IDNC schemes which grants the original packets with lower coding opportunity a higher priority to be encoded. Numerical analyses demonstrate that the lightweight IDNC schemes constructed by the new algorithm not only achieve lower completion and decoding delays in comparison with the ones constructed by the existing algorithm but also adhere closely to theoretical lower bounds, demonstrating their efficiency and practical utility.
Rina Su, Qifu Tyler Sun, Shaoteng Liu, Zhongshan Zhang, Linqi Song
IEEE Trans. Commun.4
2025 New Construction of MDS Array Codes and Explicit Characterization of Decoding Matrices
abstract
Row-Diagonal-Parity (RDP) codes and EVENODD codes are classical systematic array codes and most attention in the literature has been on the generalization of RDP codes. In this work, as generalization of not only RDP codes but also EVENODD codes, we present new construction of$\phi (L)$-dimensional$(k+r, k)$systematic array codes with$r \leq 4$, where L is an odd integer and$\phi (L)$represents the Euler’s totient function of L. We explicitly characterize sufficient conditions on the selection of L to make the codes maximum distance separable (MDS). Compared with EVENODD codes and RDP codes, the largest k that can be supported by the new codes is nearly doubled, and the asymptotic encoding complexity of the new codes is same, that is, asymptotically approaches r XORs per original data bit with increasing L and k. Moreover, for prime L,$r = 2$and$k = 2L-3$, the new code exactly achieves the optimal encoding complexity. For the case$r = 4$, the largest k that can be supported by the new codes is larger than the recently proposed so-called Variants of Extended Shortened Independent-Parity (V-ESIP) systematic array code in a number of code dimension selections, and meanwhile, the obtained explicit conditions on L to guarantee the MDS property of the new codes also apply to classical EVENODD codes and RDP codes, but are more general than well known explicit ones in the literature. The decoding process of the new array codes is also discussed. In particular, the$r\times r$block inverse matrix involved in decoding is explicitly characterized, which applies to all MDS array codes generalized from RDP or EVENODD codes in the literature.
Zhe Zhai, Qifu Tyler Sun, Shaoteng Liu, Xiangyu Chen 0004, Zongpeng Li
IEEE Trans. Commun.4
2025 Stateless and Proactive Routing for Dynamic Multicast With Deep Reinforcement Learning
abstract
Stateful multicast protocols manage multicast group memberships by maintaining state information about active groups and their members. They have seen limited adoption in the modern internet due to lack of scalability, simplicity, and flexibility. Although stateless multicast protocols, like BIER, eliminate extensive state management, they still face complex tree computation and limited scalability for concurrent requests. In this paper, we propose Hawkeye, a stateless multicast mechanism with deep reinforcement learning (DRL) for real-time responses to dynamic multicast requests with near-optimal multicast TE performance. This mechanism is suited for Software-Defined Networking (SDN) environment where the controller has a global view of the network and supports flexible configuration of network resources for traffic engineering. For real-time responses to multicast requests, we leverage DRL enhanced by a temporal convolutional network (TCN) to model the sequential feature of dynamic group membership, and thus are able to build multicast trees proactively for upcoming requests. We develop a novel source aggregation mechanism to facilitate the convergence of the DRL agent under high volume of multicast requests. Moreover, to improve the practicality and robustness of Hawkeye, we design incremental deployment and single failure handling mechanisms, which take advantages of source aggregation and fit well with multicast routing. Evaluation with real-world topologies and multicast requests demonstrates that Hawkeye responds effectively to dynamic multicast requests. Itoffers rapid routing decisions, e.g., making routing decisions in under 5ms on a tested topology, and reduces path latency variation by up to 89.5%, with less than a 10% increase in bandwidth consumption compared to the offline theoretical minimum.
Qing Li 0006, Lie Lu, Dan Zhao 0003, Zeyu Luan, Yuan Yang 0001, Yong Jiang 0001, Jingpu Duan, Ruobin Zheng, Shaoteng Liu, Dingding Chen
IEEE Trans. Netw.9
2024 Video-P2P: Video Editing with Cross-Attention Control
abstract
Video-P2P is the first framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-trained image generation models, there are currently no large-scale video generation models publicly available. Video-P2P addresses this limitation by adapting an image generation diffusion model to complete various video editing tasks. Specifically, we propose to first tune a Text-to-Set (T2S) model to complete an approximate inversion and then optimize a shared unconditional embedding to achieve accurate video inversion with a small memory cost. We further prove that it is crucial for consistent video editing. For attention control, we introduce a novel decoupled-guidance strategy, which uses different guidance strategies for the source and target prompts. The optimized unconditional embedding for the source prompt improves reconstruction ability, while an initialized unconditional embedding for the target prompt enhances editability. Incorporating the attention maps of these two branches enables detailed editing. These technical designs enable various text-driven editing applications, including word swap, prompt refinement, and attention re-weighting. Video-P2P works well on real-world videos for generating new characters while optimally preserving their original poses and scenes. It significantly outperforms previous approaches.
Shaoteng Liu, Yuechen Zhang, Wenbo Li 0002, Zhe Lin 0001, Jiaya Jia
CVPR1
2024 HGR: A Hybrid Global Graph-Based Recovery Approach for Cloud Storage Systems with Failure and Straggler Nodes
abstract
Cloud storage systems often face the issues of failure and straggler nodes. Failure is characterized as a fail-stop scenario, which refers to disk failures that can result in significant data unavailability. Straggler nodes are typically those with heavy workloads or poor performance. Usually, both failure and straggler nodes coexist, posing a significant challenge to data availability in storage systems. In such failure scenarios, parallel recovery and straggler recovery methods are commonly used as separate approaches for data recovery. However, parallel recovery methods encounter bottlenecks on the recovery path due to the presence of straggler nodes. Meanwhile, straggler recovery methods face the challenge of lacking available recovery paths in cases of multiple node failures. Scenarios involving both multiple failures and stragglers are common, yet there is a lack of efficient recovery methods for these situations. In this paper, we focus on scenarios involving video data, which occupies a significant portion of cloud storage systems, to address the above issues. We propose a Hybrid Global Graph-based Recovery (HGR) method that integrates parallel and straggler recovery approaches into a single global graph. The key idea of HGR is to construct a global graph that includes global node parameter information, enabling comprehensive coordination. We partition the global graph into two subgraphs: one containing straggler nodes and the other containing failure nodes. Resources are efficiently allocated to each subgraph to schedule recovery tasks in parallel. For data that presents significant recovery challenges, exhibits poor parallelism, has substantial tail latency, or exceeds fault tolerance limits, we employ approximate recovery methods. To demonstrate HGR's effectiveness, we conducted several experiments. The results indicate that HGR can reduce recovery time by up to 45.06% and improve I/O throughput by as much as 1.79× compared to state-of-the-art recovery methods.
Piao Hu, Huangzhen Xue, Chentao Wu, Minyi Guo, Jie Li 0002, Xiangyu Chen 0006, Shaoteng Liu, Liyang Zhou, Shenghong Xie
ICDCS7
2024 PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of Code
abstract
Text-guided diffusion models have revolutionized image generation and editing, offering exceptional realism and diversity. Specifically, in the context of diffusion-based editing, where a source image is edited according to a target prompt, the process commences by acquiring a noisy latent vector corresponding to the source image via the diffusion model. This vector is subsequently fed into separate source and target diffusion branches for editing. The accuracy of this inversion process significantly impacts the final editing outcome, influencing both essential content preservation of the source image and edit fidelity according to the target prompt. Prior inversion techniques aimed at finding a unified solution in both the source and target diffusion branches. However, our theoretical and empirical analyses reveal that disentangling these branches leads to a distinct separation of responsibilities for preserving essential content and ensuring edit fidelity. Building on this insight, we introduce “PnP Inversion,” a novel technique achieving optimal performance of both branches with just three lines of code. To assess image editing performance, we present PIE-Bench, an editing benchmark with 700 images showcasing diverse scenes and editing types, accompanied by versatile annotations and comprehensive evaluation metrics. Compared to state-of-the-art optimization-based inversion techniques, our solution not only yields superior performance across 8 editing methods but also achieves nearly an order of speed-up.
Xuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu, Qiang Xu 0001
ICLR4
2024 RL-GPT: Integrating Reinforcement Learning and Code-as-policy
abstract
Large Language Models (LLMs) have demonstrated proficiency in utilizing various tools by coding, yet they face limitations in handling intricate logic and precise control. In embodied tasks, high-level planning is amenable to direct coding, while low-level actions often necessitate task-specific refinement, such as Reinforcement Learning (RL). To seamlessly integrate both modalities, we introduce a two-level hierarchical framework, RL-GPT, comprising a slow agent and a fast agent. The slow agent analyzes actions suitable for coding, while the fast agent executes coding tasks. This decomposition effectively focuses each agent on specific tasks, proving highly efficient within our pipeline. Our approach outperforms traditional RL methods and existing GPT agents, demonstrating superior efficiency. In the Minecraft game, it rapidly obtains diamonds within a single day on an RTX3090. Additionally, it achieves SOTA performance across all designated MineDojo tasks.
Shaoteng Liu, Haoqi Yuan, Minda Hu, Yukang Chen, Shu Liu 0005, Zongqing Lu 0002, Jiaya Jia
NeurIPS1
2024 Lightweight Instantly Decodable Network Coding in Wireless Broadcast
abstract
We consider broadcasting a block of data packets to multiple users via instantly decodable network coding (IDNC) under the semi-online feedback transmission mode. In this paper, we first introduce a new class of IDNC schemes called lightweight IDNC, tailored for wireless broadcast with stringent computational load at the receiver end. Unlike traditional IDNC that may encode a larger number of original packets together, lightweight IDNC limits each coded packet to a combination of at most two original packets. We obtain lower bounds on the total completion delay that apply to arbitrary lightweight IDNC schemes. We further investigate the number of transmission rounds as another performance metric, and explicitly characterize its distribution and expectation. The characterizations apply to arbitrary partition-based IDNC schemes, including the lightweight IDNC schemes considered in this paper. A new efficient algorithm is also proposed to construct lightweight IDNC schemes which grants the original packets with lower coding opportunity a higher priority to be encoded. Numerical analyses demonstrate that the lightweight IDNC schemes constructed by the new algorithm not only achieve lower completion and decoding delays in comparison with the ones constructed by the existing algorithm but also adhere closely to theoretical lower bounds, demonstrating their efficiency and practical utility.
Rina Su, Qifu Tyler Sun, Shaoteng Liu, Zhongshan Zhang, Linqi Song
VTC Spring4
2024 MATE: When multi-agent Deep Reinforcement Learning meets Traffic Engineering in multi-domain networks
Zeyu Luan, Qing Li 0006, Yong Jiang 0001, Jingpu Duan, Ruobin Zheng, Dingding Chen, Shaoteng Liu
Comput. Networks7
2023 New Construction of (k + r,k) Systematic MDS Array Codes with r ≤ 4
abstract
Given a prime L, we present a new construction of (L−1)-dimensional (k+r,k) systematic array codes with r ≤ 4, and concretely characterize sufficient conditions on the selection of L to guarantee the codes’ MDS property. The largest possible k that can be supported by the new MDS array codes is 2L−4, nearly twice as large as that supported by classical MDS array codes such as EVENODD codes and RDP codes. Moreover, the number of XORs per original data bit required in encoding of the new codes asymptotically approaches r with increasing k and L, same as EVENODD codes and RDP codes. In addition, for the case r = 4, the explicit conditions on L we obtain to guarantee the new codes’ MDS property can also be used to guarantee the MDS property of EVENODD codes and RDP codes, but are more general than the well known ones in the literature.
Zhe Zhai, Qifu Tyler Sun, Shaoteng Liu, Xiangyu Chen 0004
ITW4
2023 Accelerating Distributed DNN Training via Transport Layer Scheduling
abstract
Communication scheduling is crucial to accelerate the training of large deep learning models, in which the transmission order of layer-wise deep neural network (DNN) tensors is determined for a better computation-communication overlap. Prior approaches adopt user-level tensor partitioning to enhance the priority scheduling with finer granularity. However, a startup time slot inserted before every tensor partition will neutralize this scheduling gain. Tuning hyper-parameters for tensor partitioning is difficult, especially when the network bandwidth is shared or time-varying in multi-tenant clusters. In this article, we propose Mercury, a simple transport layer scheduler that moves the priority scheduling to the transport layer at the packet granularity. The packets with the highest priority in the Mercury buffer will be transmitted first. Mercury achieves the near-optimal overlapping between communication and computation. It also leverages the immediate aggregation at the transport layer to enable the full overlapping of gradient push and pull. We implement Mercury in MXNet and conduct comprehensive experiments on five popular DNN models in various environments. Mercury can well adapt to dynamic communication and computation resources. Experiments show that Mercury accelerates the training by up to 130% compared to the classical PS architecture, and 104% compared to state-of-the-art tensor partitioning methods.
Qingyang Duan, Zeqin Wang, Yuedong Xu 0001, Shaoteng Liu, Jun Wu 0006, John C. S. Lui
IEEE Trans. Parallel Distributed Syst.5
2022 Mercury: A Simple Transport Layer Scheduler to Accelerate Distributed DNN Training
abstract
Communication scheduling is crucial to improve the efficiency of training large deep learning models with data parallelism, in which the transmission order of layer-wise deep neural network (DNN) tensors is determined for a better computation-communication overlap. Prior approaches adopt tensor partitioning to enhance the priority scheduling with finer granularity. However, a startup time slot inserted before each tensor partition will neutralize this scheduling gain. Tuning the optimal partition size is difficult and the application-layer solutions cannot eliminate the partitioning overhead. In this paper, we propose Mercury, a simple transport layer scheduler that does not partition the tensors, but moves the priority scheduling to the transport layer at the packet granularity. The packets with the highest priority in the Mercury buffer will be transmitted first. Mercury achieves the near-optimal overlapping between communication and computation. It leverages immediate aggregation at the transport layer to enable the coincident gradient push and parameter pull. We implement Mercury in MXNet and conduct comprehensive experiments on five DNN models in an 8-node cluster with 10Gbps Ethernet. Experimental results show that Mercury can achieve about 1.18 ~ 2.18 × speedup over vanilla MXNet, and 1.08 ~ 2.04× speedup over the state-of-the-art tensor partitioning solution.
Qingyang Duan, Zeqin Wang, Yuedong Xu 0001, Shaoteng Liu, Jun Wu 0006
INFOCOM4
2022 Blind Robust Video Watermarking Based on Adaptive Region Selection and Channel Reference
abstract
Digital watermarking technology has a wide range of applications in video distribution and copyright protection due to its excellent invisibility and convenient traceability. This paper proposes a robust blind watermarking algorithm using adaptive region selection and channel reference. By designing a combinatorial selection algorithm using texture information and feature points, the method realizes automatically selecting stable blocks which can avoid being destroyed during video encoding and complex attacks. In addition, considering human's insensitivity to some specific color components, a channel-referenced watermark embedding method is designed for less impact on video quality. Moreover, compared with other methods' embedding watermark only at low frequencies, our method tends to modify low-frequency coefficients close to mid frequencies, further ensuring stable retention of the watermark information in the video encoding process. Experimental results show that the proposed method achieves excellent video quality and high robustness against geometric attacks, compression, transcoding and camcorder recordings attacks.
Qinwei Chang, Leichao Huang, Shaoteng Liu, Hualuo Liu, Yexin Wang
ACM Multimedia3
2021 Multiple Auxiliary Networks for Single Blind Image Deblurring
abstract
Single blind image deblurring caused by a combination of multiple factors has been one of the most challenging visual tasks. Recently, many essential methods of this task are based on deep learning networks and have achieved high performance. However, most of them only apply norm pixel-wise L1-loss function as the guide of training, which is not suitable or effective enough. In this paper, we propose Multiple Auxiliary Networks (MANet) for single blind image deblurring to assist norm L1-loss function and enhance the quality of the deblurring image. The main branch of our MANet is an encoder-decoder structure made up of residual blocks, and the three auxiliary branches are the edge prediction branch, the multi-scale refinement branch, and the perceptual loss branch. The experimental results demonstrate that the proposed MANet can obtain better deblurring performance with more details than state-of-the-art methods. The code is released at github.com/ZERO2ER0/MANet.
Chen Li 0063, Qi Wang 0009, Shaoteng Liu, Xuelong Li 0001
ICASSP3
2021 Tent: Fully Test-Time Adaptation by Entropy Minimization
Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen, Trevor Darrell
ICLR3
2021 A Hybrid Approach for Detecting Prerequisite Relations in Multi-Modal Food Recipes
abstract
Modeling the structure of culinary recipes is the core of recipe representation learning. Current approaches mostly focus on extracting the workflow graph from recipes based on text descriptions. Process images, which constitute an important part of cooking recipes, has rarely been investigated in recipe structure modeling. We study this recipe structure problem from a multi-modal learning perspective, by proposing aprerequisite treeto represent recipes with cooking images at a step-level granularity. We propose a simple-yet-effective two-stage framework to automatically construct the prerequisite tree for a recipe by (1) utilizing a trained classifier to detect pairwise prerequisite relations that fuses multi-modal features as input; then (2) applying different strategies (greedy method, maximum weight, and beam search) to build the tree structure. Experiments on the MM-ReS dataset demonstrates the advantages of introducing process images for recipe structure modeling. Also, compared with neural methods which require large numbers of training data, we show that our two-stage pipeline can achieve promising results using only 400 labeled prerequisite trees as training data.
Liangming Pan, Jingjing Chen 0001, Shaoteng Liu, Chong-Wah Ngo, Min-Yen Kan, Tat-Seng Chua
IEEE Trans. Multim.3
2020 Hyperbolic Visual Embedding Learning for Zero-Shot Recognition
abstract
This paper proposes a Hyperbolic Visual Embedding Learning Network for zero-shot recognition. The network learns image embeddings in hyperbolic space, which is capable of preserving the hierarchical structure of semantic classes in low dimensions. Comparing with existing zeroshot learning approaches, the network is more robust because the embedding feature in hyperbolic space better represents class hierarchy and thereby avoid misleading resulted from unrelated siblings. Our network outperforms exiting baselines under hierarchical evaluation with an extremely challenging setting, i.e., learning only from 1,000 categories to recognize 20,841 unseen categories. While under flat evaluation, it has competitive performance as state-of-the-art methods but with five times lower embedding dimensions. Our code is publicly available.
Shaoteng Liu, Jingjing Chen 0001, Liangming Pan, Chong-Wah Ngo, Tat-Seng Chua, Yu-Gang Jiang 0001
CVPR1
2020 GREEN: a Graph REsidual rE-ranking Network for Grading Diabetic Retinopathy
Shaoteng Liu, Lijun Gong, Kai Ma 0002, Yefeng Zheng 0001
MICCAI (5)1
2020 Multi-modal Cooking Workflow Construction for Food Recipes
abstract
Understanding food recipe requires anticipating the implicit causal effects of cooking actions, such that the recipe can be converted into a graph describing the temporal workflow of the recipe. This is a non-trivial task that involves common-sense reasoning. However, existing efforts rely on hand-crafted features to extract the workflow graph from recipes due to the lack of large-scale labeled datasets. Moreover, they fail to utilize the cooking images, which constitute an important part of food recipes. In this paper, we build MM-ReS, the first large-scale dataset for cooking workflow construction, consisting of 9,850 recipes with human-labeled workflow graphs. Cooking steps are multi-modal, featuring both text instructions and cooking images. We then propose a neural encoder-decoder model that utilizes both visual and textual information to construct the cooking workflow, which achieved over 20% performance gain over existing hand-crafted baselines.
Liangming Pan, Jingjing Chen 0001, Jianlong Wu, Shaoteng Liu, Chong-Wah Ngo, Min-Yen Kan, Yu-Gang Jiang 0001, Tat-Seng Chua
ACM Multimedia4
2019 Scene Classification With Recurrent Attention of VHR Remote Sensing Images
abstract
Scene classification of remote sensing images has drawn great attention because of its wide applications. In this paper, with the guidance of the human visual system (HVS), we explore the attention mechanism and propose a novel end-to-end attention recurrent convolutional network (ARCNet) for scene classification. It can learn to focus selectively on some key regions or locations and just process them at high-level features, thereby discarding the noncritical information and promoting the classification performance. The contributions of this paper are threefold. First, we design a novel recurrent attention structure to squeeze high-level semantic and spatial features into several simplex vectors for the reduction of learning parameters. Second, an end-to-end network named ARCNet is proposed to adaptively select a series of attention regions and then to generate powerful predictions by learning to process them sequentially. Third, we construct a new data set named OPTIMAL-31, which contains more categories than popular data sets and gives researchers an extra platform to validate their algorithms. The experimental results demonstrate that our model makes great promotion in comparison with the state-of-the-art approaches.
Qi Wang 0009, Shaoteng Liu, Jocelyn Chanussot, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2018 Control under Intermittent Network Partitions
abstract
We propose a novel distributed leader election algorithm to deal with the controller and control service availability issues in programmable networks, such as Software Defined Networks (SDN) or programmable Radio Access Network (RAN). Our approach can deal with a wide range of network failures, especially intermittent network partitions, where splitting and merging of a network repeatedly occur. In contrast to traditional leader election algorithms that mainly focus on the (eventual) consensus on one leader, the proposed algorithm aims at optimizing control service availability, stability and reducing the controller state synchronization effort during intermittent network partitioning situations. To this end, we design a new framework that enables dynamic leader election based on real-time estimates acquired from statistical monitoring. With this framework, the proposed leader election algorithm has the capability of being flexibly configured to achieve different optimization objectives, while adapting to various failure patterns. Compared with two existing algorithms, our approach can significantly reduce the synchronization overhead (up to 12x) due to controller state updates, and maintain up to twice more nodes under a controller.
Shaoteng Liu, Rebecca Steinert, Dejan Kostic
ICC1
2018 Attention Based Network for Remote Sensing Scene Classification
abstract
Scene classification of very high resolution remote sensing images is becoming more and more important because of its wide range of applications. However, previous works are mainly based on handcrafted features which do not have enough adaptability and expression ability. In this paper, inspired by the attention mechanism of human visual system, we propose a novel attention based network (AttNet) for scene classification. It can focus selectively on some key areas of images so that it can abandon redundant information. Essentially, AttNet gives a way to readjust the signal of supervision, and it is one of the first successful attempts on visual attention for remote sensing scene classification. Our method is evaluated on the UC Merced Land-Use Dataset, in comparison with some state-of-the-art methods. The experimental result shows that the proposed method makes a great improvement on both convergence speed and classification accuracy, and it also shows the effectiveness of visual attention for this task.
Shaoteng Liu, Qi Wang 0009, Xuelong Li 0001
IGARSS1
2018 Flexible distributed control plane deployment
abstract
For large-scale programmable networks, flexible deployment of distributed control planes is essential for service availability and performance. However, existing approaches only focus on placing controllers whereas the consequent control traffic is often ignored. In this paper, we propose a black-box optimization framework offering the additional steps for quanti-fying the effect of the consequent control traffic when deploying a distributed control plane. Evaluating different implementations of the framework over real-world topologies shows that close to optimal solutions can be achieved. Moreover, experiments indicate that running a method for controller placement without considering the control traffic, cause excessive bandwidth usage (worst cases varying between 20.1%-50.1% more) and congestion, compared to our approach.
Shaoteng Liu, Rebecca Steinert, Dejan Kostic
NOMS1
2015 Highway in TDM NoCs
abstract
TDM (Time Division Multiplexing) is a well-known technique to provide QoS guarantees in NoCs. However, unused time slots commonly exist in TDM NoCs. In the paper, we propose a TDM highway technique which can enhance the slot utilization of TDM NoCs. A TDM highway is an express TDM connection composed of special buffer queues, called highway channels (HWCs). It can enhance the throughput and reduce data transfer delay of the connection, while keeping the quality of service (QoS) guarantee on minimum bandwidth and in-order packet delivery. We have developed a dynamic and repetitive highway setup policy which has no dependency on particular TDM NoC techniques and no overhead on traffic flows. As a result, highways can be efficiently established and utilized in various TDM NoCs.
Shaoteng Liu, Zhonghai Lu, Axel Jantsch
NOCS1
2015 MultiCS: Circuit switched NoC with multiple sub-networks and sub-channels
Shaoteng Liu, Axel Jantsch, Zhonghai Lu
J. Syst. Archit.1
2014 Parallel probe based dynamic connection setup in TDM NoCs
abstract
We propose a Time-Division Multiplexing (TDM) based connection oriented NoC with a novel double time-wheel router architecture combined with a run-time parallel probing setup method. In comparison with traditional TDM connection setup methods, our design has the following advantages: (1) it allocates paths and time slots at run-time; (2) it is fast with predictable and bounded setup latency; (3) it avoids additional resources (no auxiliary network or central processor to find and manage connections); (4) it is fully distributed and therefore it scales nicely with network size. Compared to a packet based setup method, our probe based design can reduce path setup delay by 81% and increase network load by 110% in an 8×8 mesh, while avoiding the auxiliary network. Compared to a centralized method, our solution can double the success rate, while eliminating the central resource for path setup and reducing the wire overhead. Synthesis results suggest that our design is faster and smaller than all comparable solutions.
Shaoteng Liu, Axel Jantsch, Zhonghai Lu
DATE1
2014 A Fair and Maximal Allocator for Single-Cycle On-Chip Homogeneous Resource Allocation
abstract
Traditional allocators for network-on-chip (NoC) routers suffer from either poor-matching quality or limited fairness. We propose a waterfall (WTF) allocator targeting homogeneous resource allocation, which provides single-cycle maximal matching while guaranteeing strong fairness based on the round-robin principle. It can be implemented with a loop-free structure. In 90 nm technology, the allocator operates at about 1 GHz clock frequency. We compare WTF with wave-front, separable-input-first, and separable-output-first allocators and find that it is at least 10% smaller, has 50% less delay under high load, and uses 3% less power than any of these alternatives. Also, WTF is at least as fair or clearly fairer. We also find that in a 4×4 circuit switched NoC the use of WTF gives up to 20% higher network performance.
Shaoteng Liu, Axel Jantsch, Zhonghai Lu
IEEE Trans. Very Large Scale Integr. Syst.1
2013 Analysis and Evaluation of Circuit Switched NoC and Packet Switched NoC
abstract
Circuit switched NoC has, compared to packet switching, a longer setup time, guaranteed throughput and latency, higher clock frequency, lower HW complexity, and higher energy efficiency. Depending on packet size and throughput requirements they exhibit better or worse performance. In this paper we designed a circuit switched NoC and compared that with packet switched NoC. By speculation and analysis, we propose that, as packet size increases, performance decreases for packet switched NoC, while it increases for circuit switched NoC. By close examination on the router architecture, we suggest that circuit switched NoC can operate at a higher clock frequency than packet switched NoC, and thus at zero load above a certain packet size circuit switched NoC could be better than packet switched NoC in packet delay. Experiment results support our intuitions and analysis. We find the cross-over point, above which circuit switching has lower latency, is around 30 flits/packet under low load and 60-70 flits/packet under high network load.
Shaoteng Liu, Axel Jantsch, Zhonghai Lu
DSD1
2012 Parallel probing: Dynamic and constant time setup procedure in circuit switching NoC
abstract
We propose a circuit switching Network-on-chip with a parallel probe searching setup method, which can search the entire network in constant time, only dependent on the network size but independent of the network load. Under a specific search policy, the setup procedure is guaranteed to terminate in time 3D+6 cycles, where D is the geometric distance between source and destination. If a path can be found, the method succeeds in 3D+6 cycles; if a path cannot be found, it fails in maximum 3D+6 cycles. Compared to previous work, our method can reduce the setup time and enhance the success rate of setups. Our experiments show that compared with a sequential probe searching method, this method can reduce the search time by up to 20%. Compared with a centralized channel allocator method, this method can enhance the success rate by up to 20%.
Shaoteng Liu, Axel Jantsch, Zhonghai Lu
DATE1