EDBT 2026 Demo / reviewers in the wild / expert
Ying Liu 0024
dblp:91/112-24
· DBLP profile ↗
89ranked-venue papers
6as first author
42since 2021 · last 2026
0000-0002-4919-1130ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 53 · 30 since 2021Security and privacy · 13 · 6 since 2021Systems, architecture and hardware · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Understanding the IPv6 Address Usage Strategies of Top Internet Services
Lin He 0004, Zedong Jia, Daguo Cheng, Jinlong E, Yuhan Du, Guanglei Song, Ying Liu 0024, Xingang Shi, Shenglin Zhang, Jiahai Yang 0001, Mingwei Xu 0001 |
ICC | 7 |
| 2026 | Divide, Predict, Conquer: Adaptive Internet-wide Service Discovery with Limited Seeds
Daguo Cheng, Zedong Jia, Ying Liu 0024, Lin He 0004, Le Gai, Jiuzhou Zhang, Chentian Wei, Zhaoan Wang, Jinlong E |
INFOCOM | 3 |
| 2026 | SpecNet-Agent: Network-Aware Speculation Control for QoS in Agentic Generative AI Services
Le Gai, Lin He 0004, Chentian Wei, Zedong Jia, Daguo Cheng, Ying Liu 0024 |
IWQoS | 6 |
| 2026 | SwitchTAD: Defending deep learning-based website fingerprinting attacks with programmable switches
Lin He 0004, Xiaoyi Shi, Yifan Yang 0009, Jinlong E, Ying Liu 0024 |
Comput. Networks | 6 |
| 2026 | AddrProbe: An Internet-Wide Active IPv6 Address Probing System With Limited SeedsabstractWith the large-scale deployment of IPv6, it is becoming more and more important to probe active IPv6 addresses on the global Internet. However, the vast address space and the random distribution of active addresses make the probing process full of challenges, especially for the probing of IPv6 prefixes without seed addresses. Furthermore, the widespread existence of IPv6 aliased prefixes also causes significant trouble for probing. In this paper, we presentAddrProbe, an active IPv6 address probing system, which dynamically probes all global routing prefixes based on learned fine-grained address patterns from limited seed addresses and quickly detects aliased prefixes during probing. The evaluation results show thatAddrProbeachieves a hit rate of 23%-45% with all routing prefixes announced by the BGP system, which is 6.6-13× that of current state-of-the-art approaches (no more than 4%). Moreover, we find 1.2×1033aliased addresses characterized by the detected aliased prefixes, covering 6,412 routing prefixes, which is a 107× and 5.9× improvement over existing methods, respectively. Finally, an IPv6 Hitlist is constructed based on the long-term probing results, which contains 562M addresses covering 190K routing prefixes and 29K ASes. These widely distributed addresses are meaningful for analyzing IPv6 address assignments and some other IPv6 measurement activities. Daguo Cheng, Lin He 0004, Qilei Yin, Guangxing Han, Boran Jin, Ying Liu 0024, Guanglei Song, Jinlong E, Tiankai Yang 0001, Jiahai Yang 0001 |
IEEE Trans. Netw. | 7 |
| 2026 | Do Not Fall Into the Trap: Efficiently Discovering IPv6 Fully Responsive Prefixes in the Wild
Lin He 0004, Chentian Wei, Daguo Cheng, Qilei Yin, Boran Jin, Zhaoan Wang, Xiaoteng Pan, Sixu Zhou, Ying Liu 0024, Shenglin Zhang, Fuchao Tan, Wenmao Liu |
IEEE Trans. Netw. | 9 |
| 2025 | TopoMiner: Efficient IPv6 Topology DiscoveryabstractTopology discovery can be used to obtain the connectivity and operational status of network devices by discovered interfaces. By performing topology discovery on networks, administrators can better understand and manage networks and improve network reliability and security. However, the large IPv6 address space, sparse address distribution, and unknown address assignment policies make it infeasible to simply and roughly probe the IPv6 network topology. To this end, we design an efficient IPv6 interface-level topology discovery method called TopoMiner. It performs multiple rounds of probing the IPv6 topology with the given set of IPv6 prefixes and locates the high-density interface regions based on the results of each round of topology discovery. TopoMiner then performs prefix expansion and address generation for the high-density interface address regions. At the same time, TopoMiner also uses some of the prefixes that have been probed to regenerate the target addresses. In this way, TopoMiner can dynamically adjust the breadth and depth of topology discovery and thus complete multiple rounds of topology discovery within the packetsending budget. In real-word probing, TopoMiner achieves a$3 \sim 7 \times$enhancement in probing efficiency compared to state-of-the-art topology discovery methods. Hongwei Li 0021, Lin He 0004, Guanglei Song, Daguo Cheng, Jiahai Yang 0001, Ying Liu 0024 |
ICC | 8 |
| 2025 | Gungnir: Autoregressive Model for Unified Generation of IPv6 Fully Responsive PrefixesabstractWith the widespread adoption of IPv6, its vast address space presents significant challenges for network asset discovery. Traditional exhaustive scanning approaches are no longer practical, while strategies based on target generation algorithms are increasingly undermined by the existence of Fully Responsive Prefixes (FRPs)—prefixes in which all addresses appear responsive to probing. FRPs distort scanning results, introducing bias and inefficiency. Existing FRP probing techniques suffer from limited scalability, poor accuracy, and restricted applicability in large-scale IPv6 environments.To this end, we propose Gungnir, a multi-protocol unified FRP probing algorithm based on autoregressive semantic modeling. Gungnir captures the intricate relationships between FRP patterns and their influencing factors through a deep semantic learning architecture. It leverages prefix inference and a granularity correction mechanism to accurately predict and validate FRPs, while mitigating errors from incorrect prefix-length estimation. Extensive experiments demonstrate that Gungnir outperforms state-of-the-art techniques, achieving up to 27× higher efficiency, 4.2× wider address space coverage, and broader coverage of both autonomous systems and routing prefixes under the same probing budget. Beyond performance, we further analyze the service and port distributions of the discovered FRPs, uncovering operational patterns and potential security implications. These insights offer valuable guidance for IPv6 measurement, address discovery, and network defense. Chentian Wei, Ying Liu 0024, Lin He 0004, Daguo Cheng |
ICNP | 2 |
| 2025 | Lightning in the Dark: Uncovering Global IPv6 Router Interfaces and Their Security ImplicationsabstractThe IPv6 routing infrastructure is an important part of the modern Internet, and the collection of its interface addresses is greatly significant in network security, performance optimization, and measurement analysis. However, existing methods suffer from two major problems: the lack of flexibility in budget allocation across probing rounds and the absence of a dynamic hop limit adjustment mechanism based on feedback. These problems lead to the low hit rate and inefficiency of existing methods for discovering router interfaces, which seriously hinders the comprehensive knowledge of IPv6 routing infrastructure.To this end, we propose Helixir, a feedback-based, high hit-rate, and efficient IPv6 router interface discovery system. Helixir’s core design includes a dynamic budget allocation mechanism across probing rounds, an inter-prefix budget allocation strategy that adequately trades off exploration and exploitation, and a hop limit selection method based on Thompson sampling. Real-world experiments show that with a 100M budget, Helixir achieves a hit rate 3.64× that of state-of-the-art methods on the BGP prefixes dataset, and Helixir successfully discovers over 31 million IPv6 router interface addresses in total within half an hour. In addition, a systematic security analysis of the discovered router interfaces shows that many devices open sensitive ports and expose hundreds of potential CVE vulnerabilities, highlighting the security risks in the IPv6 network. Ying Liu 0024, Lin He 0004, Xiaoyi Shi, Yifan Yang 0009, Chentian Wei, Daguo Cheng, Jiahai Yang 0001 |
ICNP | 2 |
| 2025 | SubRecon: Efficient Internet-Wide IPv6 Subnet Discovery and Its ApplicationsabstractThe vastness of the IPv6 address space has led to the common practice of allocating prefixes to end users rather than individual addresses. Users can assign these prefixes as a single subnet or divide them into multiple subnets for different purposes. Allocation strategies vary significantly in terms of prefix granularity, and identifying the actual granularity of subnet assignments is crucial for improving measurement efficiency, accuracy, and for better IPv6 network management. However, no existing method can discover IPv6 subnets at an Internet-Wide scale.To this end, we propose SubRecon, an Internet-Wide IPv6 subnet discovery system. SubRecon consists of two key phases: subnet delimitation and target expansion. In the subnet delimitation phase, we perform a systematic scan across the entire IPv6 address space without relying on any existing seed dataset. This phase adopts a top-down approach, probing prefixes from the shortest to the longest in a hierarchical manner. At each level, we recursively refine prefixes and discard sub-prefixes that do not meet the convergence condition. This pruning strategy eliminates redundant probes in unallocated regions, significantly reducing the search space and improving probing efficiency. To further improve coverage, the target expansion phase leverages the active address dataset as an auxiliary input. It identifies active addresses not covered by previously discovered subnets, expands them into new candidate target prefixes, and performs another round of subnet delimitation. This helps enhance the completeness and coverage of the final discovered subnet set. Experimental results show that SubRecon discovers 8,381,974 IPv6 subnets across 14,147 autonomous systems, and the resulting subnet list can serve as high-quality input for topology discovery. Additionally, during the subnet discovery process, SubRecon identifies a large number of last-hop router interfaces, discovering approximately 10 million more than the current state-of-the-art methods. Ying Liu 0024, Lin He 0004, Yifan Yang 0009, Xiaoyi Shi, Daguo Cheng, Chentian Wei, Yun Fan, Guanglei Song |
ICNP | 2 |
| 2025 | Poster: TopoHunter: Enabling Efficient and High-Coverage Active IPv6 Topology DiscoveryabstractWe introduce TopoHunter, an efficient IPv6 Internet topology discovery system. The central concept of TopoHunter is to allocate more probing resources to target prefix spaces that yield greater topological benefits, as well as to their surrounding areas. To achieve this, we design a feedback-based target generation module comprised of a Target Prefix Probing Value Forest that maintains the estimated probing values of hierarchical target prefix spaces. Our system has successfully discovered the most extensive and complete IPv6 topology map to date, comprising over 144 million router interfaces and 251 million edges, covering 72.83% of autonomous systems and 43.36% of routing prefixes announced by the BGP system. Lin He 0004, Hongwei Li 0021, Guanglei Song, Wentong Wang, Daguo Cheng, Enhuan Dong, Chenglong Li 0006, Hui Zhang 0141, Jinlong E, Ying Liu 0024, Jiahai Yang 0001 |
IMC | 12 |
| 2025 | 6Map: Enabling Fast Active IPv6 Address Discovery with Programmable Switches
Lin He 0004, Yifan Yang 0009, Xiaoyi Shi, Daguo Cheng, Jinlong E, Ying Liu 0024, Dong Zhang 0010 |
INFOCOM | 7 |
| 2025 | Bringing New Life to Old Tools: Measuring Source Address Validation Deployment with 6in4 TunnelsabstractSource Address Validation (SAV) is a security mechanism deployed at network boundaries to prevent packets with illegal source addresses from crossing these boundaries. While SAV plays an important role in mitigating source address spoofing, its deployment across the global Internet remains limited. Measuring SAV deployment is essential for enhancing the understanding of the security landscape of networks. In this study, we propose a novel method for measuring SAV deployment based on 6 in 4 tunnels, complementing existing measurement work. Using this method, we measure inbound SAV for IPv4 and outbound SAV for IPv6, obtaining results from 12,417 and 2,104 Autonomous Systems (ASes), respectively. Based on our measurements, we analyze factors that may influence SAV deployment, including network address space size, AS type, and geographical location. Additionally, we provide a global heatmap of spoofable rates for networks in different countries and regions. Note that the measuring method using bin4 tunnels we introduce is not only applicable to SAV measurements but also holds potential for other measurement tasks, such as connectivity testing and transmission path discovery. This method offers a new way for large-scale measurement tasks across different networks, which may benefit future research. Jiaxing Guo, Lin He 0004, Daguo Cheng, Xingang Shi, Ying Liu 0024 |
IWQoS | 5 |
| 2025 | APCC: Enabling Reliable IPv6 Covert Communication with Aliased PrefixesabstractCovert communication ensures undetectable information exchange between parties while posing risks when exploited for malicious purposes. Existing network covert channels face challenges in reliability, throughput, and stealthiness due to packet loss, limited capacity, and detectable anomalies. This paper proposes APCC, a reliable covert communication system that leverages IPv6 aliased prefixes for the first time, where secret data is embedded in the Interface Identifier field of IPv6 addresses. APCC enhances stealthiness through encryption and camouflage strategies (e.g., traffic blending and rate control). Additionally, it employs a reliable transmission mechanism incorporating sequence numbering, acknowledgment, and retransmission to mitigate packet loss and reordering. It ensures deployment flexibility across ICMPv6, UDP, or TCP protocols. Evaluations in real-world and simulated environments demonstrate that APCC achieves 100% accuracy under high latency ($\mathbf{8 0 0 ~ m s ~ R T T}$) and packet loss (10%), with throughput up to 34.7 Kbps. Detection tests show APCC evades major intrusion detection systems (Snort, Zeek) except for Suricata's TCP alerts. We also propose mitigation measures to reduce APCC's potential negative impact. This work highlights critical vulnerabilities in IPv6 infrastructure while advancing robust covert communication methodologies for high-stakes scenarios. Zhaoan Wang, Lin He 0004, Daguo Cheng, Ying Liu 0024 |
IWQoS | 4 |
| 2025 | SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingabstractThe performance of collective communication schedules is crucial for the efficiency of machine learning jobs and GPU cluster utilization. Existing open-source collective communication libraries (such as NCCL and RCCL) rely on fixed schedules and cannot adjust to varying topology and model requirements. State-of-the-art collective schedule synthesizers (such as TECCL and TACCL) utilize Mixed Integer Linear Program for modeling but encounter search space explosion and scalability challenges. In this paper, we propose SyCCL, a scalable collective schedule synthesizer that aims to synthesize near-optimal schedules in tens of minutes for production-scale machine-learning jobs. SyCCL leverages collective and topology symmetries to decompose the original collective communication demand into smaller sub-demands within smaller topology subsets. SyCCL proposes efficient search strategies to quickly explore potential sub-demands, synthesizes corresponding sub-schedules, and integrates these sub-schedules into complete schedules. Our 32-A100 testbed and production-scale simulation experiments show that SyCCL improves collective performance by up to 127% while reducing synthesis time by 2 to 4 orders of magnitude compared to state-of-the-art efforts. Jiamin Cao, Shangfeng Shi, Weisen Liu, Yifan Yang 0009, Yichi Xu, Zhilong Zheng, Yu Guan 0005, Kun Qian 0021, Ying Liu 0024, Mingwei Xu 0001, Ning Wang 0001, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 10 |
| 2025 | TGW: Operating an Efficient and Resilient Cloud Gateway at Scale
Yifan Yang 0009, Lin He 0004, Xiaoyi Shi, Yichi Xu, Jinlong E, Ying Liu 0024, Zhuang Yuan, Hengyang Xu |
USENIX ATC | 8 |
| 2025 | Miresga: Accelerating Layer-7 Load Balancing with Programmable SwitchesabstractAs online cloud services expand rapidly, layer-7 load balancing has become indispensable for maintaining service availability and performance. The emergence of programmable switches with both high performance and a certain degree of flexibility has made it possible to apply programmable switches to load balancing. Nevertheless, the limited memory capacity and the relatively sluggish speed of table entry insertion and deletion of programmable switches have severely constrained their performance. Xiaoyi Shi, Lin He 0004, Yifan Yang 0009, Ying Liu 0024 |
WWW | 5 |
| 2024 | ChatScam: Unveiling the Rising Impact of ChatGPT on Domain Name AbuseabstractSince 2022, ChatGPT has been a big breakthrough in technology, creating lots of discussions online. It has had big effects in different areas, but in cybersecurity, it is both good and bad. There has been a lot of misuse, especially with squatting domains. Our research aims to understand this misuse and the potential threats it poses. We develop a novel method that looks at historical Passive DNS (PDNS) data. Based on the two-stage identification, our method can efficiently and accurately collect ChatGPT-related squatting domains. In the end, we found over 1.3 million ChatGPT-related squatting domains, part of which were shared with the security community. Our findings show that these squatting domains are increasing quickly. This is the case whether the keywords related to ChatG PT are registered with the domain registrar or set up on sub domains. Even though the number of domains is increasing, only 5.3 % set up meaningful content on their websites. After digging into their web contents, we found that these web sites show various signs of misuse, such as promotion on illegal underground websites and emerging fraudulent activities related to dialogue features. The security community is not fully aware of these threats yet. We are the first to conduct a large-scale quantitative analysis of ChatGPT-related abusive behavior. We believe that our work unveils the abuse ecosystem surrounding ChatGPT-related squatting domains. We hope to underscore the urgent need for increased attention and protective measures against ChatGPT-related domain abuse. Mingxuan Liu 0006, Zhenglong Jin, Jiahai Yang 0001, Baoiun Liu, Hai-Xin Duan, Ying Liu 0024, Ximeng Liu, Shujun Tang |
DSN | 6 |
| 2024 | Luori: Active Probing and Evaluation of Internet-Wide IPv6 Fully Responsive PrefixesabstractWith the large-scale deployment and application of IPv6, IPv6 network measurements will become increasingly important. However, a special type of IPv6 prefix called Fully Responsive Prefix (FRP) is having a significant impact on IPv6 measurement campaigns, which is defined as all addresses under a prefix responding to scans. Obviously, there cannot be a real responder behind each of these addresses. To reveal the current status and impact of Internet-wide IPv6 FRPs, we propose for the first time an active probing method for Internet-wide IPv6 FRPs, Luori, which transforms the active probing process under IPv6 huge prefix space (potential range of prefix presence) into a dynamic search process in a tree based on reinforcement learning, achieving efficient probing of arbitrary routing prefixes. The evaluation results show that Luori found 31.7K largest FRPs in a single Internet-wide probing with 11 M budget, covering$1.5 \times 10^{30}$address space, which is$10^{6} \times$that of existing methods. More importantly, after six months of Internet-wide probing, we have found 516 K largest FRPs, which covers$1.3 \times 10^{33}$address space and 795 ASes, making it the largest publicly known FRP list. Based on this list, we screen out$20 \%$of the addresses covered by FRPs from a well-known IPv6 active address dataset. Furthermore, we further analyze and find that the distribution of these FRPs is extensive and their implementation methods are diverse, which can provide beneficial references for the practical application of FRPs. We also make this list publicly available and maintain it long-term for use and study by relevant researchers. Daguo Cheng, Lin He 0004, Chentian Wei, Qilei Yin, Boran Jin, Zhaoan Wang, Xiaoteng Pan, Sixu Zhou, Ying Liu 0024, Shenglin Zhang, Fuchao Tan, Wenmao Liu |
ICNP | 9 |
| 2024 | Overlooked Backdoors: Investigating 6to4 Tunnel Nodes and Their Exploitation in the WildabstractAs native IPv6 adoption increases, the use of 6to4 tunnels has declined, yet they remain a significant security concern in today’s Internet. This study investigates the real-world deployment of 6to4 tunnels, revealing their current scale, characteristics, and security implications. We identify open 6to4 relays in 216 countries and 13,114 autonomous systems, noting stable short-term counts but a long-term decline. We analyze the security of these nodes and find over 578k nodes vulnerable to address spoofing and packet injection. Additionally, we present several under-emphasized scenarios where open 6to4 nodes are abused, including leveraging services on 6to4 nodes as traffic amplifiers, circumventing restrictions using multiple 6to4 addresses, and connecting 6to4 nodes to render attacks untraceable. Jiaxing Guo, Lin He 0004, Ying Liu 0024 |
IPCCC | 3 |
| 2024 | P4runpro: Enabling Runtime Programmability for RMT Programmable SwitchesabstractProgrammable switches have revolutionized network operations by enabling the flexible customization of packet processing logic using language like P4. However, changing the programs running on the switch requires disturbing traffic and suspending other unrelated programs. In this paper, we present P4runpro, enabling runtime data plane updates with dynamic resource allocation. The P4runpro data plane abstracts hardware resources and defines dynamically reconfigurable atomic operations that form packet processing logic. P4runpro provides runtime programming interfaces called P4runpro primitives for the operator to write high-level programs. We have designed the P4runpro compiler to automatically and consistently link the P4runpro programs to the running data plane. We implement our prototype on a Tofino switch. We implement 15 example runtime programs using P4runpro to demonstrate its generality and expressiveness. Our evaluation results show that compared to the state-of-the-art, P4runpro can respond within hundreds of milliseconds, achieve an average of 60% to 80% dynamic resource utilization, concurrently run ≈0.6K to ≈2.8K programs, and introduce lower overhead. Our case studies illustrate the benefit of runtime programming and prove the same functionality between P4runpro and conventional P4 programs. Yifan Yang 0009, Lin He 0004, Xiaoyi Shi, Jiamin Cao, Ying Liu 0024 |
SIGCOMM | 6 |
| 2024 | Investigating Deployment Issues of DNS Root Server Instances From a China-Wide ViewabstractDNS root servers are the starting point of most DNS queries. To ensure their security and stability, multiple anycast instances are operated worldwide, and new root instances have been rapidly deployed in recent years. Apart from authorized instances managed by Root Server System, some networks equip unauthorized instances to hijack queries from clients. Despite various root instances handling queries within their residing networks, few studies have focused on the deployment issues of these instances. In this paper, we provide the first study to reveal the deployment issues of root instances from a nationwide view. With the support of 7,860 vantage points, we utilized a suite of methodologies to identify the deployment of unauthorized instances. 54 vantage points witnessed the evidence of unauthorized instances, and 70.4% of them further observed security issues of unauthorized instances, including DoS, unavailability of DNSSEC validation, and vulnerable DNS software. Additionally, we utilized the side-channel information of censorship mechanisms to measure the catchment area of authorized instances. We found that most authorized instances in the Chinese mainland serve with limited catchment areas due to restricted BGP policies. Through discussions with ISPs and network operators, we make recommendations to improve the deployment status of different root instances. Fenglu Zhang, Baojun Liu 0002, Chaoyi Lu, Yunpeng Xing, Hai-Xin Duan, Ying Liu 0024, Liyuan Chang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2023 | Silence is not Golden: Disrupting the Load Balancing of Authoritative DNS ServersabstractAuthoritative nameservers are delegated to provide the final resource record. Since the security and robustness of DNS are critical to the general operation of the Internet, domain name owners are required to deploy multiple candidate nameservers for traffic load balancing. Once the load balancing mechanism is compromised, an adversary can manipulate a large number of legitimate DNS requests to a specified candidate nameserver. As a result, it may not only bypass the defense mechanisms used to filter malicious traffic that can overload the victim nameserver, but also lowers the bar for DNS traffic hijacking and cache poisoning attacks. Fenglu Zhang, Baojun Liu 0002, Eihal Alowaisheq, Jianjun Chen 0005, Chaoyi Lu, Linjian Song, Ying Liu 0024, Hai-Xin Duan, Min Yang 0002 |
CCS | 8 |
| 2023 | PTStore: Lightweight Architectural Support for Page Table IsolationabstractPage tables are critical data structures in kernels, serving as the trust base of most mitigation solutions. Their integrity is thus crucial but is often taken for granted. Existing page table protection solutions usually provide insufficient security guarantees, require heavy hardware, or introduce high overheads. In this paper, we present a novel lightweight hardware-software co-design solution, PTStore, consisting of a secure region storing page tables and tokens verifying page table pointers. Evaluation results on FPGA-based prototypes show that PTStore only introduces <0.92% hardware overheads and <0.86% performance overheads, but provides strong security guarantees, showing that PTStore is efficient and effective. Wende Tan, Yangyu Chen 0002, Yuan Li 0061, Ying Liu 0024, Chao Zhang 0008 |
DAC | 4 |
| 2023 | HyperClassifier: Accurate, Extensible and Scalable Traffic Classification with Programmable SwitchesabstractTraffic classification provides substantial benefits for service differentiation, security policy enforcement, and traffic engineering. However, accurately classifying large volumes of network traffic using existing solutions is pretty challenging, as they are typically implemented on commodity servers with slow CPUs for packet processing. To address this, we leverage the opportunity provided by emerging programmable switches and propose HyperClassifier as a solution to achieve accurate, extensible, and scalable traffic classification. HyperClassifier designs an efficient classifying table with an effective flow expiration mechanism that enables lightweight packet inspection on resource-limited switches. We implement an open-source prototype of HyperClassifier on a hardware Tofino switch and conduct extensive evaluations. The results of our evaluation demonstrate that, compared to existing solutions, HyperClassifier can provide orders of magnitude higher classification throughput with comparable classification accuracy. Yichi Xu, Jiamin Cao, Menghao Zhang 0001, Ying Liu 0024, Mingwei Xu 0001 |
ICC | 5 |
| 2023 | Wolf in Sheep's Clothing: Evaluating Security Risks of the Undelegated Record on DNS Hosting ServicesabstractLeveraging DNS for covert communications is appealing since most networks allow DNS traffic, especially the ones directed toward renowned DNS hosting services. Unfortunately, most DNS hosting services overlook domain ownership verification, enabling miscreants to host undelegated DNS records of a domain they do not own. Consequently, miscreants can conduct covert communication through such undelegated records for whitelisted domains on reputable hosting providers. In this paper, we shed light on the emerging threat posed by undelegated records and demonstrate their exploitation in the wild. To the best of our knowledge, this security risk has not been studied before. Fenglu Zhang, Baojun Liu 0002, Eihal Alowaisheq, Lingyun Ying, Xiang Li 0108, Zaifeng Zhang, Ying Liu 0024, Hai-Xin Duan, Min Zhang 0054 |
IMC | 8 |
| 2023 | LogSummary: Unstructured Log Summarization for Software SystemsabstractWe propose LogSummary, an automatic, unsupervised end-to-end log summarization framework for software system maintenance in this work. LogSummary obtains the summarized triples of necessary logs for a given log sequence. It integrates a novel information extraction method that considers semantic information and domain knowledge with a new triple-ranking approach using the global knowledge learned from all logs. Given the lack of a publicly-available gold standard for log summarization, we have manually labeled the summaries of four open-source log datasets and made them publicly available. The evaluation of these datasets and the case studies on real-world logs demonstrate that LogSummary produces highly representative (average ROUGE F1 score of 0.741) summaries efficiently. We have packaged LogSummary into an open-source toolkit and hope it can be a standard baseline and benefit future log summarization works. Weibin Meng, Federico Zaiter, Ying Liu 0024, Shenglin Zhang, Shimin Tao, Yichen Zhu 0001, En Wang, Dan Pei |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2023 | Bolt: Scalable and Cost-Efficient Multistring Pattern Matching With Programmable SwitchesabstractMulti-string pattern matching is a crucial building block for many network security applications and thus of great importance. Since every byte of a packet has to be inspected by a large set of patterns, it often becomes a bottleneck of these applications and dominates the performance of an entire system. Many existing studies have been devoted to alleviating this performance bottleneck either by algorithm optimization or hardware acceleration. However, neither one provides the desired scalability and costs that keep pace with the drastic increase in network bandwidth and traffic today. To address these issues, in this paper, we present BOLT, a scalable and cost-efficient multi-string pattern matching system leveraging the capability of emerging programmable switches. BOLT combines the following techniques: (1) an efficient state encoding scheme to fit a large number of strings into the limited memory on a programmable switch; (2) a variable$k$-stride transition mechanism to increase the throughput significantly with the same level of memory cost; and(3)a compactpattern2rulemapping method to accommodate multiple co-existing strings in one rule. We implement a prototype of BOLT and make its source code publicly available. Extensive evaluations demonstrate that BOLT can provide multi-hundred Gbps throughput and scales well with various pattern sets and workloads. Menghao Zhang 0001, Chang Liu 0021, Ying Liu 0024, Mingwei Xu 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2022 | PACMem: Enforcing Spatial and Temporal Memory Safety via ARM Pointer AuthenticationabstractMemory safety is a key security property that stops memory corruption vulnerabilities. Different types of memory safety enforcement solutions have been proposed and adopted by sanitizers or mitigations to catch and stop such bugs, at the development or deployment phase. However, existing solutions either provide partial memory safety or have overwhelmingly high performance overheads. Yuan Li 0061, Wende Tan, Zhizheng Lv, Songtao Yang 0001, Mathias Payer, Ying Liu 0024, Chao Zhang 0008 |
CCS | 6 |
| 2022 | Measuring the Practical Effect of DNS Root Server Instances: A China-Wide Case Study
Fenglu Zhang, Chaoyi Lu, Baojun Liu 0002, Hai-Xin Duan, Ying Liu 0024 |
PAM | 5 |
| 2022 | Firebolt: Finding Bugs in Programmable Data Plane Generators
Jiamin Cao, Yu Zhou 0008, Chen Sun 0005, Lin He 0004, Zhaowei Xi, Ying Liu 0024 |
USENIX ATC | 6 |
| 2022 | TurboNet: Faithfully Emulating Networks With Programmable SwitchesabstractFaithfully emulating networks is critical for verifying the correctness and effectiveness of new networking-related designs. Existing network experiment platforms either cannot faithfully emulate the functionality and performance of production networks or cannot scale well due to cost constraints. In this paper, we proposeTurboNet, a new network emulator that utilizes one or more programmable switches to achieve faithful emulation of the network data plane and control plane. For data plane emulation, we propose a series of key designs, such as port mapper, queue mapper, and delayed queue, to emulate network topologies and performance metrics with high flexibility and accuracy. For control plane emulation, we support static routing configurations, distributed routing agents, and the centralized routing controllers. Meanwhile, we provide APIs for operators to simplify network emulation tasks. We implementTurboNeton Tofino switches. Evaluation results show that: (1) On the data plane,TurboNetcan flexibly emulate various topologies, such as an 8-ary fat-tree with only one programmable switch and a 10-ary fat-tree with four programmable switches; (2) On the control plane,TurboNetsupports about 200 BGP agents on a single programmable switch with a CPU usage of 25%; (3)TurboNetcan accurately emulate different network performance metrics such as 10−8link loss, and microsecond to millisecond link delay. Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Lin He 0004, Mingwei Xu 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | CoFilter: High-Performance Switch-Accelerated Stateful Packet Filter for Bare-Metal ServersabstractAs one of the most critical cloud services, Bare-Metal Servers (BMS) introduce stringent performance requirements on data center networks (DCN). Stateful packet filter is an integral DCN component of ensuring connection security for BMS. However, the off-the-shelf stateful packet filters either are costly for cloud DCNs or introduce significant performance bottlenecks. In this article, we presentCoFilter, which leverages low-cost programmable switches to accelerate the stateful packet filter for BMS.CoFilteruses (1)stateful process partitionto enable complex stateful packet filtering logic on programmability-limited switching ASICs, (2)state compressionto track tens of millions of connections with constrained hardware memory, and (3)per-tenant packet rate limit and tenant-aware flow migrationto achieve efficient performance isolation among different tenants. Overall,CoFilterimplements a high-performance stateful packet filter via the co-design of programmable switching ASIC and CPU. We evaluateCoFilterunder various data center traffic traces with real-world flow distributions. The evaluation results show thatCoFilterremarkably outperforms NetFilter, i.e., forwarding packets at line rate (13x throughput of NetFilter), keeping packet delay within 1us, and freeing a significant quantity of CPU cores, with rather small memory usage, i.e., accommodating over$10^7$connections with only 16MB SRAM. Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Lin He 0004, Chen Sun 0005, Yangyang Wang 0001, Mingwei Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | ROLoad: Securing Sensitive Operations with Pointee IntegrityabstractSensitive operations (e.g. control-flow transfers) are attractive targets for attackers. To protect them from being hijacked, we propose a new solution ROLoad to guarantee the integrity of their operands, which are loaded from (potentially corrupted) memory. We extend the RISC-V instruction set, implement an FPGA-based prototype of ROLoad, and then demonstrate two specific defense applications. Results show that this solution only costs few extra hardware resources (< 3.32%). However, it could enable many lightweight (e.g. with overheads less than 0.31%) defenses, and provide broader and stronger security guarantees than existing hardware solutions, e.g. ARM BTI and Intel CET. Wende Tan, Yuan Li 0061, Chao Zhang 0008, Xingman Chen, Songtao Yang 0001, Ying Liu 0024 |
DAC | 6 |
| 2021 | Making Multi-String Pattern Matching Scalable and Cost-Efficient with Programmable Switching ASICsabstractMulti-string pattern matching is a crucial building block for many network security applications, and thus of great importance. Since every byte of a packet has to be inspected by a large set of patterns, it often becomes a bottleneck of these applications and dominates the performance of an entire system. Many existing works have been devoted to alleviate this performance bottleneck either by algorithm optimization or hardware acceleration. However, neither one provides the desired scalability and costs that keep pace with the dramatic increase of the network bandwidth and network traffic today. In this paper, we present BOLT, a scalable and cost-efficient multi-string pattern matching system leveraging the capability of emerging programmable switches. BOLT combines the following two techniques, a smart state encoding scheme to fit a large number of strings into the limited memory on the programmable switch, and a variable k-stride transition mechanism to increase the throughput significantly with the same level of memory costs. We implement a prototype of BOLT and make its source code publicly available. Extensive evaluations demonstrate that BOLT could provide orders of magnitude improvement in throughput which is scalable with pattern sets and workloads, and could also significantly decrease the number of entries and memory requirement. Menghao Zhang 0001, Chang Liu 0021, Ying Liu 0024, Xuya Jia, Mingwei Xu 0001 |
INFOCOM | 5 |
| 2021 | pSAV: A Practical and Decentralized Inter-AS Source Address Validation Service FrameworkabstractSource IP address spoofing has been a major vulnerability of the Internet for many years. Although much work has been done to study the problem extensively, spoofing continues to occur frequently and has led to many serious network attacks. Inter-AS source address validation (SAV) is considered an important defense method for AS to filter spoofed packets. However, existing work has been unable to drive inter-AS SAV deployment into practice due to the lack of deployment incentives and trust foundation.In this paper, we propose a practical and decentralized inter-AS SAV service framework, pSAV, to promote inter-AS SAV deployment. pSAV increases deployment incentives by treating SAV as a payable service and dividing the participant ASes into service subscribers, providers, and auditors. On the control plane, pSAV leverages blockchain as a trust foundation to provide service subscriptions and audits with automatic incentive allocation. On the data plane, pSAV leverages P4-programmable switches to provide flexible and high-performance SAV services. We prototype the pSAV control plane based on Hyperledger Fabric and implement various SAV techniques on Barefoot Tofino switches. The evaluation results show that (1) on the control plane, pSAV blockchain can provide high-performance service transactions (hundreds of transactions per second with second latency), and (2) on the data plane, pSAV can provide various high-throughput (hundreds of Gbps) SAV services using only one programmable switch. Jiamin Cao, Ying Liu 0024, Mingxing Liu, Lin He 0004, Yihao Jia |
IWQoS | 2 |
| 2021 | Towards Chain-Aware Scaling Detection in NFV with Reinforcement LearningabstractElastic scaling enables dynamic and efficient re-source provisioning in Network Function Virtualization (NFV) to serve fluctuating network traffic. Scaling detection determines the appropriate time when a virtual network function (VNF) needs to be scaled, and its precision and agility profoundly affect system performance. Previous heuristics define fixed control rules based on a simplified or inaccurate understanding of deployment environments and workloads. Therefore, they fail to achieve optimal performance across a broad set of network conditions.In this paper, we propose a chain-aware scaling detection mechanism, namely CASD, which learns policies directly from experience using reinforcement learning (RL) techniques. Furthermore, CASD incorporates chain information into control policies to efficiently plan the scaling sequence of VNFs within a service function chain. This paper makes the following two key technical contributions. Firstly, we develop chain-aware representations, which embed global chains of arbitrary sizes and shapes into a set of embedding vectors based on graph embedding techniques. Secondly, we design an RL-based neural network model to make scaling decisions based on chain-aware representations. We implement a prototype of CASD, and its evaluation results demonstrate that CASD reduces the overall system cost and improves system performance over other baseline algorithms across different workloads and chains. Lin He 0004, Lishan Li, Ying Liu 0024 |
IWQoS | 3 |
| 2021 | TAP: A Traffic-Aware Probabilistic Packet Marking for Collaborative DDoS MitigationabstractIn recent years, Distributed Denial-of-Service (DDoS) attacks have become more rampant and continue to be one of the most serious security threats facing network infrastructure. In a classic DDoS attack, the attacker controls numerous bots from many sources to send a significant volume of traffic to flood the victim end or the bottleneck link. In practical networks, it is inefficient and costly to request all partner routers to collaboratively mitigate DDoS attacks. The common feature of DDoS attacks is the abnormal distribution of traffic to the victim. In this paper, we propose TAP, a collaborative DDoS mitigation framework, based on traffic-aware probabilistic packet marking (PPM). TAP enables the victim to select a few hit routers as collaborators to mitigate attack traffic efficiently depending on the traffic distribution. Our evaluation results show that TAP greatly reduces attack traffic within seconds and mitigate the damage caused by DDoS with less overhead, which demonstrates that TAP is an effective, efficient, and rapid-response scheme for collaborative DDoS mitigation. Mingxing Liu, Ying Liu 0024, Ke Xu 0002, Lin He 0004, Xiaoliang Wang 0004, Yangfei Guo, Weiyu Jiang |
MSN | 2 |
| 2021 | From WHOIS to WHOWAS: A Large-Scale Measurement Study of Domain Registration Privacy under the GDPR
Chaoyi Lu, Baojun Liu 0002, Yiming Zhang 0009, Zhou Li 0001, Fenglu Zhang, Hai-Xin Duan, Ying Liu 0024, Joann Qiongna Chen, Jinjin Liang, Zaifeng Zhang, Shuang Hao 0001, Min Yang 0002 |
NDSS | 7 |
| 2021 | Towards securing Duplicate Address Detection using P4
Lin He 0004, Peng Kuang, Ying Liu 0024, Gang Ren 0003, Jiahai Yang 0001 |
Comput. Networks | 3 |
| 2021 | LogClass: Anomalous Log Identification and Classification With Partial LabelsabstractLogs are imperative in the management process of networks and services. However, manually identifying and classifying anomalous logs is time-consuming, error-prone, and labor-intensive. Additionally, rule-based approaches cannot tackle the challenges underlying anomalous log identification and classification resulting from new types of logs and partial labels. We propose LogClass, a framework to automatically and robustly identify and classify anomalous logs for network and service based onpartial labels. LogClass combines a word representation method, a positive and unlabeled learning (PU learning) model, and a machine learning classifier. Besides, we propose a novel Inverse Location Frequency (ILF) method to weight the words of logs in feature construction properly. We evaluate the performance of LogClass based on 18 million+ real-world switch logs and six public log datasets. It achieves 99.56% and 98% F1 scores in anomalous log identification on switch logs and publicly available supercomputer logs, respectively, and very-close-to-one F1 score in anomalous log classification. Moreover, we have conducted extensive experiments to demonstrate LogClass’ superior performance in addressing partial labels and new types of logs. Weibin Meng, Ying Liu 0024, Shenglin Zhang, Federico Zaiter, Zhaoyang Yu 0002, Dan Pei |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2021 | PAVI: Bootstrapping Accountability and Privacy to IPv6 InternetabstractAccountability and privacy are considered valuable but conflicting properties in the Internet, which at present does not provide native support for either. Past efforts to balance accountability and privacy in the Internet have unsatisfactory deployability due to the introduction of new communication identifiers, and because of large-scale modifications to fully deployed infrastructures and protocols. The IPv6 is being deployed around the world and this trend will accelerate. In this paper, we propose a private and accountable proposal based on IPv6 called PAVI that seeks to bootstrap accountability and privacy to the IPv6 Internet without introducing new communication identifiers and large-scale modifications to the deployed base. A dedicated quantitative analysis shows that the proposed PAVI achieves satisfactory levels of accountability and privacy. The results of the evaluation of a PAVI prototype show that it incurs little performance overhead, and is widely deployable. Lin He 0004, Gang Ren 0003, Ying Liu 0024, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2020 | Lies in the Air: Characterizing Fake-base-station Spam Ecosystem in ChinaabstractFake base station (FBS) has been exploited by criminals to attack mobile users by spamming fraudulent messages for over a decade. Despite that prior work has proposed several techniques to mitigate this issue, FBS spam is still a long-standing challenging issue in some countries, such as China, and causes billions of dollars of financial loss every year. Therefore, understanding and exploring the thematic strategies in the FBS spam ecosystem at a large scale would improve the defense mechanisms. Yiming Zhang 0009, Baojun Liu 0002, Chaoyi Lu, Zhou Li 0001, Hai-Xin Duan, Shuang Hao 0001, Mingxuan Liu 0006, Ying Liu 0024 |
CCS | 8 |
| 2020 | Finding Cracks in Shields: On the Security of Control Flow Integrity MechanismsabstractControl-flow integrity (CFI) is a promising technique to mitigate control-flow hijacking attacks. In the past decade, dozens of CFI mechanisms have been proposed by researchers. Despite the claims made by themselves, the security promises of these mechanisms have not been carefully evaluated, and thus are questionable. Yuan Li 0061, Chao Zhang 0008, Xingman Chen, Songtao Yang 0001, Ying Liu 0024 |
CCS | 6 |
| 2020 | P4DAD: Securing Duplicate Address Detection Using P4abstractDuplicate Address Detection (DAD) is an essential part of the Neighbor Discovery Protocol (NDP), which determines whether the IPv6 address of a node conflicts with those of other nodes. Due to the lack of verification of NDP messages, DAD is vulnerable to DoS attacks. Existing solutions suffer from high complexity, need to modify the NDP, or have a single point of failure.To solve the above problems, we propose P4DAD, a secure DAD mechanism described by P4. By creating and maintaining binding entries between IPv6 address and link-layer property of host, P4DAD can filter spoofed NDP messages in an in-network manner to prevent DoS attacks on DAD without modifications to the NDP or host stack. We implement a prototype of P4DAD and evaluate it in terms of functionality, performance, and scalability. Evaluation results show that P4DAD can prevent DoS attacks on DAD successfully with negligible overhead, and has satisfactory scalability. Peng Kuang, Ying Liu 0024, Lin He 0004 |
ICC | 2 |
| 2020 | NetView: Towards On-Demand Network-Wide Telemetry in the Data CenterabstractNetwork telemetry is to collect information (e.g., hop latency, throughput) from network devices. Network-wide telemetry is critical for operators to understand the quality of network performance and to diagnose on-going failures. The state-of-the-art telemetry approaches are far from ideal as they are unable to fully satisfy diverse requirements of operators, specifically for on-demand, full coverage, and scalable telemetry. In this paper, we provide a new framework of network telemetry for data center networks, called NetView. NetView can support various telemetry applications and frequencies on demand, monitoring each device via proactively sending dedicated probes. Technically, NetView leverages source routing to forward probes, achieving full coverage. Besides, a series of probe generation algorithms largely reduce probe number, providing high scalability. The evaluation shows that NetView reduces the bandwidth occupancy by more than two orders of magnitude compared with Pingmesh and INT-path, and conducts network-wide telemetry for large-scale data center network using only one vantage server, without bringing about resources bottleneck. Yunsenxiao Lin, Yu Zhou 0008, Zhengzheng Liu, Yangyang Wang 0001, Mingwei Xu 0001, Jun Bi, Ying Liu 0024 |
ICC | 8 |
| 2020 | A Semantic-aware Representation Framework for Online Log AnalysisabstractLogs are one of the most valuable data sources for large-scale service management. Log representation, which converts unstructured texts to structured vectors or matrices, serves as the the first step towards automated log analysis. However, the current log representation methods neither represent domain-specific semantic information of logs, nor handle the out-of-vocabulary (OOV) words of new types of logs at runtime. We propose Log2Vec, a semantic-aware representation framework for log analysis. Log2Vec combines a log-specific word embedding method to accurately extract the semantic information of logs, with an OOV word processor to embed OOV words into vectors at runtime. We present an analysis on the impact of OOV words and evaluate the performance of the OOV word processor. The evaluation experiments on four public production log datasets demonstrate that Log2Vec not only fixes the issue presented by OOV words, but also significantly improves the performance of two popular log-based service management tasks, including log classification and anomaly detection. We have packaged Log2Vec into an open-source toolkit and hope that it can be used for future research. Weibin Meng, Ying Liu 0024, Shenglin Zhang, Federico Zaiter, Bingjin Chen, Dan Pei |
ICCCN | 2 |
| 2020 | LogParse: Making Log Parsing Adaptive through Word ClassificationabstractLogs are one of the most valuable data sources for large-scale service (e.g., social network, search engine) maintenance. Log parsing serves as the the first step towards automated log analysis. However, the current log parsing methods are not adaptive. Without intra-service adaptiveness, log parsing cannot handle software/firmware upgrade because learned templates cannot match new type of logs. In addition, without cross-service adaptiveness, the logs of a new type of service cannot be accurately parsed when this service is newly deployed. We propose LogParse, an adaptive log parsing framework, to support intra-service and cross-service incremental template learning and update. LogParse turns the template generation problem into a word classification problem and learns the features of template words and variable words. We evaluate LogParse on four public production log datasets. The results demonstrate that LogParse supports accurate adaptive template update (increased from 0.559 to nearly 1.0 parsing accuracy), and a trained LogParse is adaptive for a brand new service’s log parsing. Because of LogParse’s adaptiveness, we also apply LogParse to an interesting application, log compression and deployed log compression in a top cloud service provider. We package LogParse into an open-source toolkit. Weibin Meng, Ying Liu 0024, Federico Zaiter, Shenglin Zhang, Yichen Zhu 0001, En Wang, Shimin Tao, Dian Yang, Dan Pei |
ICCCN | 2 |
| 2020 | TurboNet: Faithfully Emulating Networks with Programmable SwitchesabstractFaithfully emulating networks is critical for verifying the correctness and effectiveness of new networking-related designs. Existing network experiment platforms either cannot faithfully emulate functionality and performance of production networks or cannot scale well because of cost limitations. In this paper, we propose TurboNet, a new network emulator that leverages one programmable switch to enable faithful emulation of both network data plane and control plane. For data plane emulation, we present a series of key designs such as port mapper, queue mapper, and delayed queue to emulate network topologies and performance metrics with high flexibility and accuracy. For control plane emulation, we support static routing configurations, distributed routing agents, and the centralized routing controller. Meanwhile, we provide API for operators to simplify network emulation tasks. We implement TurboNet on a Tofino switch. The evaluation results show that: (1) TurboNet can flexibly emulate various topologies such as the 8-ary fat-tree on the data plane and support about 200 BGP agents with 25% CPU usage on the control plane; (2) TurboNet can accurately emulate different network performance metrics, including 400Gbps linerate background traffic injection, as small as 10-8link loss, and microsecond-level to millisecond-level link delay. Jiamin Cao, Yu Zhou 0008, Ying Liu 0024, Mingwei Xu 0001, Yongkai Zhou |
ICNP | 3 |
| 2020 | CDN Judo: Breaking the CDN DoS Protection with Itself
Run Guo, Baojun Liu 0002, Shuang Hao 0001, Jia Zhang 0004, Hai-Xin Duan, Kaiwen Shen, Jianjun Chen 0005, Ying Liu 0024 |
NDSS | 9 |
| 2020 | NetView: Towards on-demand network-wide telemetry in the data center
Yunsenxiao Lin, Yu Zhou 0008, Zhengzheng Liu, Yangyang Wang 0001, Mingwei Xu 0001, Jun Bi, Ying Liu 0024 |
Comput. Networks | 8 |
| 2019 | Multi-modal Representation Learning for Successive POI RecommendationabstractSuccessive POI recommendation is a fundamental problem for location-based social networks (LBSNs). POI recommendation takes a variety of POI context information (e.g. spatial location and textual comment) and user preference into consideration. Existing POI recommendation systems mainly focus on part of the POI context and user preference with a specific modeling, which loses valuable information from other aspects. In this paper, we propose to construct a multi-modal check-in graph, a heterogeneous graph that combines five check-in aspects in a unified way. We further propose a multi-modal representation learning model based on the graph to jointly learn POI and user representations. Finally, we employ an attentional recurrent neural network based on the representations for successive POI recommendation. Experiments on a public dataset studies the effects of modeling different aspects of check-in records and demonstrates the effectiveness of the method in improving POI recommendation performance. Lishan Li, Ying Liu 0024, Lin He 0004, Gang Ren 0003 |
ACML | 2 |
| 2019 | NFVMP: An Architecture for NFV Applications from Multiple ProvidersabstractWith the evolution of Network Function Virtualization (NFV), small- and medium-size enterprises are increasingly outsourcing their network functions (NFs) to the cloud. In order to obtain cost-effective and high-performance services, customers may utilize NFs from various cloud providers. Current NFV architectures are mostly designed on the assumption of one provider, which utilize centralized controllers for orchestration. In multi-provider scenario, autonomy of each provider leads to infeasibility of a centralized controller. In this paper, we propose NFVMP, a novel NFV architecture that enables cooperation of providers to compose intended chains, while maintaining dependency of each provider. NFVMP adopts a fresh view that regards the sub-chain of each provider as first-class entity, and supports sub-chain composition and flexibility. We implement a prototype, which demonstrates that NFVMP supports correct chain deployment across providers with a negligible performance overhead on latency and throughput. Lishan Li, Ying Liu 0024, Gang Ren 0003 |
APNOMS | 2 |
| 2019 | TraffickStop: Detecting and Measuring Illicit Traffic Monetization Through Large-Scale DNS AnalysisabstractIllicit traffic monetization is a type of Internet fraud that hijacks users' web requests and reroutes them to a traffic network (e.g., advertising network), in order to unethically gain monetary rewards. Despite its popularity among Internet fraudsters, our understanding of the problem is still limited. Since the behavior is highly dynamic (can happen at any place including client-side, transport-layer and server-side) and selective (could target a regional network), prior approaches like active probing can only reveal a small piece of the entire ecosystem. So far, questions including how this fraud works at a global scale and what fraudsters' preferred methods are, still remain unanswered. To fill the missing pieces, we developed TraffickStop the first system that can detect this fraud passively. Our key contribution is a novel algorithm that works on large-scale DNS logs and efficiently discovers abnormal domain correlations. TraffickStop enables the first landscape study of this fraud, and we have some interesting findings. By analyzing over 231 billion DNS logs of two weeks, we discovered 1,457 fraud sites. Regarding its scale, the fraud sites receive more than 53 billion DNS requests within one year, and a company could lose up to 53K dollars per day due to fraud traffic. We also discovered two new strategies that are leveraged by fraudsters to evade inspection. Our work provides new insights into illicit traffic monetization, raises its public awareness, and contributes to a better understanding and ultimate elimination of this threat. Baojun Liu 0002, Zhou Li 0001, Peiyuan Zong, Chaoyi Lu, Hai-Xin Duan, Ying Liu 0024, Sumayah A. Alrwais, XiaoFeng Wang 0001, Shuang Hao 0001, Yaoqi Jia, Yiming Zhang 0009, Kai Chen 0012, Zaifeng Zhang |
EuroS&P | 6 |
| 2019 | Address Protection-as-a-Service an Inter-AS Framework for IP Spoofing ResilienceabstractIP spoofing, which is generally used for anonymity and amplification, constantly leads to pervasive distributed denial-of-service (DDoS) attacks. To mitigate IP spoofing, source address validation is divided into access network, intra-autonomous system (AS), and inter-AS levels. However, because of ambiguous incentives, heterogeneous demands, and fragile trust, techniques for the inter-AS level fail in practice, and thus, IP spoofing is still considered as an almost open vulnerability of the entire Internet. In this study, we aim to transform the inter-AS source address validation into an "address protection" service, and we mitigate IP spoofing through an economics-driven framework - apf ('a'ddress 'p'rotection 'f'ramework). In such a protection, the addresses belonging to one AS can be prevented from being spoofed by others. Behind the framework, such a service will be consolidated by a unified trust anchor with a uniform interface, and deployer ASes will be free to select their preferred techniques and invoke the service when needed. Based on the empirical data and theoretical analysis, we prove that the service is acceptable for triggering economics-driven implementation under the guidance of the apf framework. Yihao Jia, Ying Liu 0024, Gang Ren 0003 |
GLOBECOM | 2 |
| 2019 | GSDM: Graph-Based Scaling Detection Model in Network Function VirtualizationabstractNetwork function virtualization (NFV) is an emerging technology, which aims at replacing proprietary network function hardware devices with software- based network function (NF) applications. In NFV, it is critical to conduct dynamic and elastic resource allocation in accordance with varying workloads. State- of-the-art scaling detection algorithms are designed based on either traffic rate or runtime status. However, they cannot make precise and agile scaling decisions due to diversity of input traffic and noisy measurements. In this paper, we propose a novel graph-based scaling detection model (GSDM) using deep learning techniques. GSDM is a neural network model, which selects scaling action based on feature sequences of each NF and its adjacent NFs. Neural network provides a scalable way to incorporate both traffic rate and runtime status into control policy. Furthermore, adjacent NFs' feature information is also injected. As a result, GSDM empirically learns policies that achieves precise and agile scaling detections and improves system performance. Our implemented prototype demonstrates that GSDM achieves superior performance in precision and recall, and outperforms other schemes in a variety of network conditions. Lishan Li, Ying Liu 0024, Gang Ren 0003 |
GLOBECOM | 2 |
| 2019 | CoFilter: A High-Performance Switch-Accelerated Stateful Packet Filter for Bare-Metal ServersabstractAs one of the most critical cloud services, Bare-metal Servers introduce stringent performance requirements on data center networks (DCN). Stateful packet filter is an integral DCN component of ensuring connection security for bare-metal servers. However, the off-the-shelf hardware-based and software-based stateful packet filters either are prohibitively costly for cloud DCNs or introduce significant performance bottlenecks. In this paper, we present CoFilter, which employs cheap programmable switches to accelerate the stateful packet filter for bare-metal servers. CoFilter consists of two key designs. First, to support complex stateful packet filtering logic in programmability-limited switching ASICs, CoFilter partitions the stateful packet filtering logic between programmable ASICs and switch CPU. Most packets are directly processed in switching ASICs to achieve high performance, while only a small number of packets go to switch CPU for connection tracking. Second, to track massive connections with constrained hardware memory, CoFilter employs hash to compress connection states and provides an efficient settlement for hash collisions. We build a prototype of CoFilter and evaluate it on the Tofino switch under various data center traffic traces with real-world flow distribution. The evaluation shows that CoFilter largely outperforms NetFilter, i.e., forwarding packets at line rate (13x throughput of NetFilter), keeping packet delay at 1us, and freeing a significant quantity of CPU cores. Furthermore, CoFilter presents great scalability and accommodates over ten million connections with only 16MB SRAM. Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Chen Sun 0005, Yangyang Wang 0001, Jun Bi |
ICCCN | 2 |
| 2019 | LogAnomaly: Unsupervised Detection of Sequential and Quantitative Anomalies in Unstructured LogsabstractRecording runtime status via logs is common for almost every computer system, and detecting anomalies in logs is crucial for timely identifying malfunctions of systems. However, manually detecting anomalies for logs is time-consuming, error-prone, and infeasible. Existing automatic log anomaly detection approaches, using indexes rather than semantics of log templates, tend to cause false alarms. In this work, we propose LogAnomaly, a framework to model unstructured a log stream as a natural language sequence. Empowered by template2vec, a novel, simple yet effective method to extract the semantic information hidden in log templates, LogAnomaly can detect both sequential and quantitive log anomalies simultaneously, which were not done by any previous work. Moreover, LogAnomaly can avoid the false alarms caused by the newly appearing log templates between periodic model retrainings. Our evaluation on two public production log datasets show that LogAnomaly outperforms existing log-based anomaly detection methods. Weibin Meng, Ying Liu 0024, Yichen Zhu 0001, Shenglin Zhang, Dan Pei, Shimin Tao |
IJCAI | 2 |
| 2019 | An End-to-End, Large-Scale Measurement of DNS-over-Encryption: How Far Have We Come?abstractDNS packets are designed to travel in unencrypted form through the Internet based on its initial standard. Recent discoveries show that real-world adversaries are actively exploiting this design vulnerability to compromise Internet users' security and privacy. To mitigate such threats, several protocols have been proposed to encrypt DNS queries between DNS clients and servers, which we jointly term as DNS-over-Encryption. While some proposals have been standardized and are gaining strong support from the industry, little has been done to understand their status from the view of global users. Chaoyi Lu, Baojun Liu 0002, Zhou Li 0001, Shuang Hao 0001, Hai-Xin Duan, Mingming Zhang 0010, Chunying Leng, Ying Liu 0024, Zaifeng Zhang |
Internet Measurement Conference | 8 |
| 2019 | Bootstrapping Accountability and Privacy to IPv6 Internet without Starting from ScratchabstractAccountability and privacy are considered valuable but conflicting properties in the Internet, which at present does not provide native support for either. Past efforts to balance accountability and privacy in the Internet have unsatisfactory deployability due to the introduction of new communication identifiers, and because of large-scale modifications to fully deployed infrastructures and protocols. The IPv6 is being deployed around the world and this trend will accelerate. In this paper, we propose a private and accountable proposal based on IPv6 called PAVI that seeks to bootstrap accountability and privacy to the IPv6 Internet without introducing new communication identifiers and large-scale modifications to the deployed base. A dedicated quantitative analysis shows that the proposed PAVI achieves satisfactory levels of accountability and privacy. The results of evaluation of a PAVI prototype show that it incurs little performance overhead, and is widely deployable. Lin He 0004, Gang Ren 0003, Ying Liu 0024 |
INFOCOM | 3 |
| 2019 | Resident Evil: Understanding Residential IP Proxy as a Dark ServiceabstractAn emerging Internet business is residential proxy (RESIP) as a service, in which a provider utilizes the hosts within residential networks (in contrast to those running in a datacenter) to relay their customers' traffic, in an attempt to avoid server- side blocking and detection. With the prominent roles the services could play in the underground business world, little has been done to understand whether they are indeed involved in Cybercrimes and how they operate, due to the challenges in identifying their RESIPs, not to mention any in-depth analysis on them. In this paper, we report the first study on RESIPs, which sheds light on the behaviors and the ecosystem of these elusive gray services. Our research employed an infiltration framework, including our clients for RESIP services and the servers they visited, to detect 6 million RESIP IPs across 230+ countries and 52K+ ISPs. The observed addresses were analyzed and the hosts behind them were further fingerprinted using a new profiling system. Our effort led to several surprising findings about the RESIP services unknown before. Surprisingly, despite the providers' claim that the proxy hosts are willingly joined, many proxies run on likely compromised hosts including IoT devices. Through cross-matching the hosts we discovered and labeled PUP (potentially unwanted programs) logs provided by a leading IT company, we uncovered various illicit operations RESIP hosts performed, including illegal promotion, Fast fluxing, phishing, malware hosting, and others. We also reverse engi- neered RESIP services' internal infrastructures, uncovered their potential rebranding and reselling behaviors. Our research takes the first step toward understanding this new Internet service, contributing to the effective control of their security risks. Xianghang Mi, Xuan Feng 0005, Xiaojing Liao, Baojun Liu 0002, XiaoFeng Wang 0001, Feng Qian 0001, Zhou Li 0001, Sumayah A. Alrwais, Limin Sun 0001, Ying Liu 0024 |
IEEE Symposium on Security and Privacy | 10 |
| 2018 | A Reexamination of Internationalized Domain Names: The Good, the Bad and the UglyabstractInternationalized Domain Names (IDNs) are domain names containing non-ASCII characters. Despite its installation in DNS for more than 15 years, little has been done to understand how this initiative was developed and its security implications. In this work, we aim to fill this gap by studying the IDN ecosystem and cyber-attacks abusing IDN. In particular, we performed by far the most comprehensive measurement study using IDNs discovered from 56 TLD zone files. Through correlating data from auxiliary sources like WHOIS, passive DNS and URL blacklists, we gained many insights. Our discoveries are multi-faceted. On one hand, 1.4 million IDNs were actively registered under over 700 registrars, and regions within east Asia have seen prominent development in IDN registration. On the other hand, most of the registrations were opportunistic: they are currently not associated with meaningful websites and they have severe configuration issues (e.g., shared SSL certificates). What is more concerning is the rising trend of IDN abuse. So far, more than 6K IDNs were determined as malicious by URL blacklists and we also identified 1,516 and 1,497 IDNs showing high visual and semantic similarity to reputable brand domains (e.g., apple.com). Meanwhile, brand owners have only registered a few of these domains. Our study suggests the development of IDN needs to be re-examined. New solutions and proposals are needed to address issues like its inadequate usage and new attack surfaces. Baojun Liu 0002, Chaoyi Lu, Zhou Li 0001, Ying Liu 0024, Hai-Xin Duan, Shuang Hao 0001, Zaifeng Zhang |
DSN | 4 |
| 2018 | Rapid Deployment of Anomaly Detection Models for Large Number of Emerging KPI StreamsabstractInternet-based services monitor and detect anomalies on KPIs (Key Performance Indicators, say CPU utilization, number of queries per second, response latency) of their applications and systems in order to keep their services reliable. This paper identifies a common, important, yet little-studied problem of KPI anomaly detection: rapid deployment of anomaly detection models for large number of emerging KPI streams, without manual algorithm selection, parameter tuning, or new anomaly labeling for any newly emerging KPI streams. We propose the first framework ADS (Anomaly Detection through Self-training) that tackles the above problem, via clustering and semi-supervised learning. Our extensive experiments using real-world data show that, with the labels of only the 5 cluster centroids of 70 historical KPI streams, ADS achieves an averaged best F-score of 0.92 on 81 new KPI streams, almost the same as a state-of-art supervised approach, and greatly outperforming a state-of-art unsupervised approach by 61.40% on average. Jiahao Bu, Ying Liu 0024, Shenglin Zhang, Weibin Meng, Qitong Liu, Xiaotian Zhu, Dan Pei |
IPCCC | 2 |
| 2018 | Device-Agnostic Log Anomaly Classification with Partial LabelsabstractAnomaly classification, i.e., detecting whether a network device is anomalous and determining its anomaly category if yes, plays a crucial role in troubleshooting. Compared to KPI curves, device logs contain too much more valuable information for anomaly classification. However, the regular expression based anomaly classification techniques cannot tackle the challenges lying in log anomaly classification. We propose LogClass, a data-driven framework to detect and classify anomalies based on device logs. LogClass combines a word representation method and the PU learning model to construct device-agnostic vocabulary with partial labels. We evaluate LogClass on tens of millions of switch logs collected from several real-world datacenters owned by a top global search engine. Our results show that LogClass achieves 99.515% F1 score in anomalous log detection, 95.32% Macro-F1 and 99.74% Micro-F1 in anomalous log classification in a computationally efficient manner. Weibin Meng, Ying Liu 0024, Shenglin Zhang, Dan Pei, Xulong Luo |
IWQoS | 2 |
| 2018 | Who Is Answering My Queries: Understanding and Characterizing Interception of the DNS Resolution Path
Baojun Liu 0002, Chaoyi Lu, Hai-Xin Duan, Ying Liu 0024, Zhou Li 0001, Shuang Hao 0001, Min Yang 0002 |
USENIX Security Symposium | 4 |
| 2018 | Unsupervised Anomaly Detection via Variational Auto-Encoder for Seasonal KPIs in Web ApplicationsabstractTo ensure undisrupted business, large Internet companies need to closely monitor various KPIs (e.g., Page Views, number of online users, and number of orders) of its Web applications, to accurately detect anomalies and trigger timely troubleshooting/mitigation. However, anomaly detection for these seasonal KPIs with various patterns and data quality has been a great challenge, especially without labels. In this paper, we proposed Donut, an unsupervised anomaly detection algorithm based on VAE. Thanks to a few of our key techniques, Donut greatly outperforms a state-of-arts supervised ensemble approach and a baseline VAE approach, and its best F-scores range from 0.75 to 0.9 for the studied KPIs from a top global Internet company. We come up with a novel KDE interpretation of reconstruction for Donut, making it the first VAE-based anomaly detection algorithm with solid theoretical explanation. Wenxiao Chen, Nengwen Zhao, Zeyan Li 0001, Jiahao Bu, Zhihan Li 0002, Ying Liu 0024, Youjian Zhao, Dan Pei, Zhaogang Wang, Honglin Qiao |
WWW | 7 |
| 2018 | GAGMS: a requirement-driven general address generation and management system
Ying Liu 0024, Lin He 0004, Gang Ren 0003 |
Sci. China Inf. Sci. | 1 |
| 2018 | FUNNEL: Assessing Software Changes in Web-Based ServicesabstractThe detection of performance changes in software change roll-outs in Internet-based services is crucial for an operations team, because it allows timely roll-back of a software change when performance degrades unexpectedly. However, it is infeasible to manually investigate millions of performance measurements of many roll-outs. In this paper, we present an automated tool, FUNNEL, for rapid and robust impact assessment of software changes in large Internet-based services. FUNNEL automatically collects the related performance measurements for each software change. To detect significant performance behavior changes, FUNNEL adopts singular spectrum transform (SST) algorithm as the core algorithm, uses various techniques to improve its robustness and reduce its computational cost, and applies a difference-in-difference (DiD) method to differentiate the true causality from the random correlations between the performance change and the software change. Evaluation through historical data in real-word services shows that FUNNEL achieves accuracy of more than 99.7 percent. Compared with previous methods, FUNNEL's detection delay is 38.02 to 64.99 percent shorter, and its computation speed is 4.59-7,098 times faster. In real deployment, FUNNEL achieves a 98.21 percent precision, high robustness, fast detection speed, and shows its capability in detecting unexpected behavior changes. Shenglin Zhang, Ying Liu 0024, Dan Pei, Xianping Qu, Shimin Tao, Zhi Zang, Xiaowei Jing, Mei Feng |
IEEE Trans. Serv. Comput. | 2 |
| 2017 | Revisiting inter-AS IP spoofing let the protection drive source address validationabstractIP spoofing, which is prevalently used for anonymity and reflection attacks, has shown increasing destructive power in recent years. Although certain source address validation solutions have been standardized by the Internet Engineering Task Force, few networks are willing to adopt them in view of the deficiency of deployment benefits. Actually, all the source address validation solutions face the problem of a lack of deployability. In this paper, we summarize the key points describing deployability and propose a new security service-inter-autonomous-system (AS) Source Address Protection (iSAP). Technically, by increasing the possibility of keeping the source address belonging to one AS from being the victim of reflection flooding, iSAP improves the deployers ability to prevent IP spoofing and increases incremental deployability. In reality, such a service can also be regarded as a new profit opportunity for ASes and it could progress gradually once it is well commercialized. Based on simulations with real Internet topology data, the results illustrate that iSAP can protect ASes from being reflected with only a few deployers, exhibiting a high potential to mitigate reflection flooding with modest resource consumption. Yihao Jia, Ying Liu 0024, Gang Ren 0003, Lin He 0004 |
IPCCC | 2 |
| 2017 | Syslog processing for switch failure diagnosis and prediction in datacenter networksabstractSyslogs on switches are a rich source of information for both post-mortem diagnosis and proactive prediction of switch failures in a datacenter network. However, such information can be effectively extracted only through proper processing of syslogs, e.g., using suitable machine learning techniques. A common approach to syslog processing is to extract (i.e., build) templates from historical syslog messages and then match syslog messages to these templates. However, existing template extraction techniques either have low accuracies in learning the “correct” set of templates, or does not support incremental learning in the sense the entire set of templates has to be rebuilt (from processing all historical syslog messages again) when a new template is to be added, which is prohibitively expensive computationally if used for a large datacenter network. To address these two problems, we propose a frequent template tree (FT-tree) model in which frequent combinations of (syslog) words are identified and then used as message templates. FTtree empirically extracts message templates more accurately than existing approaches, and naturally supports incremental learning. To compare the performance of FT-tree and three other template learning techniques, we experimented them on two-years' worth of failure tickets and syslogs collected from switches deployed across 10+ datacenters of a tier-1 cloud service provider. The experiments demonstrated that FT-tree improved the estimation/prediction accuracy (as measured by F1) by 155% to 188%, and the computational efficiency by 117 to 730 times. Shenglin Zhang, Weibin Meng, Jiahao Bu, Sen Yang 0001, Ying Liu 0024, Dan Pei, Jun (Jim) Xu, Xianping Qu |
IWQoS | 5 |
| 2016 | CCDN: Content-Centric Data Center NetworksabstractData center networks continually seek higher network performance to meet the ever increasing application demand. Recently, researchers are exploring the method to enhance the data center network performance by intelligent caching and increasing the access points for hot data chunks. Motivated by this, we come up with a simple yet useful caching mechanism for generic data centers, i.e., a server caches a data chunk after an application on it reads the chunk from the file system, and then uses the cached chunk to serve subsequent chunk requests from nearby servers. To turn the basic idea above into a practical system and address the challenges behind it, we design content-centric data center networks (CCDNs), which exploits an innovative combination of content-based forwarding and location [Internet Protocol (IP)]-based forwarding in switches, to correctly locate the target server for a data chunk on a fully distributed basis. Furthermore, CCDN enhances traditional content-based forwarding to determine the nearest target server, and enhances traditional location (IP)-based forwarding to make high utilization of the precious memory space in switches. Extensive simulations based on real-world workloads and experiments on a test bed built with NetFPGA prototypes show that, even with a small portion of the server's storage as cache (e.g., 3%) and with a modest content forwarding information base size (e.g., 1000 entries) in switches, CCDN can improve the average throughput to get data chunks by 43% compared with a pure Hadoop File System (HDFS) system in a real data center. Dan Li 0001, Fangxin Wang 0001, Anke Li, K. K. Ramakrishnan, Ying Liu 0024, Xue (Steve) Liu |
IEEE/ACM Trans. Netw. | 6 |
| 2015 | Rapid and robust impact assessment of software changes in large internet-based servicesabstractThe detection of performance changes in software change roll-outs in Internet-based services is crucial for an operations team, because it allows timely roll-back of a software change when performance degrades unexpectedly. However, it is infeasible to manually investigate millions of performance measurements of many roll-outs. Shenglin Zhang, Ying Liu 0024, Dan Pei, Xianping Qu, Shimin Tao, Zhi Zang |
CoNEXT | 2 |
| 2015 | MIFO: Multi-path Interdomain ForwardingabstractToday's interdomain routing is traffic agnostic when determining the single, best forwarding path. Naturally, as it does not adapt to congestion, the path chosen is not always optimal. In this paper, we focus on designing a multi-path interdomain forwarding (MIFO) mechanism, where AS border routers adaptively forward outbound traffic from a congested default path to an alternative path, without touching the interdomain routing protocols. Different from previous efforts which enable multi-path on control plane, MIFO achieves multi-path on data plane. The multiple alternative forwarding paths are obtained by exploring local BGP RIB. Multi-path forwarding on data plane can create a loop even within a stable network. MIFO solves this problem with a simple and practical approach. Several other challenges are also addressed including preventing cycling packet between iBGP peers and choosing the best alternative path from among multiple candidates. Our evaluations show that MIFO significantly improves the end-to-end throughput at the AS level, compared to traditional BGP and MIRO. For example, with only 50% of the ASes being MIFO capable, a significant percentage of the flows (about 40%) can use at least 50% of the inter-AS link capacity. In contrast, BGP and MIRO routing make less effective use of the inter-AS links, with only 7% and 17% of the flows can be so. Finally, we have developed a prototype implementation of MIFO on Linux with the forwarding engine in the kernel, with the routing daemon developed on XORP platform. The experiments on a test bed built with prototypes show that MIFO can improves the aggregate throughput by 81% compared with BGP routing. Dan Li 0001, Ying Liu 0024, Dan Pei, K. K. Ramakrishnan |
ICPP | 3 |
| 2015 | An inter-AS path vector filter: towards elimination of false negativesabstractIP spoofing based attacks remains a serious and open security problem due to the fact that the current Internet implements no source address authentication mechanisms. A series of anti-spoofing practices have long been proposed while their actual implementation seems far from satisfactory. Route based filters were extensively studied in the design of Inter-AS source address validation methods. Traditional route based filters only use route direction information to establish filtering rules, causing inherited fake negatives. A novel inter-AS filter based on route path vector is proposed to reduce or even eliminate such fake negatives in this article. We name the filter IPVF (Inter-AS Path Vector Filter), which utilizes the route information of both path and distance, exhibits measurable increase in performance and incurs acceptable additional bandwidth cost. Moreover, traditional route based filtering rules is easy to be deduced by attackers. Since the filtering rules of IPVF could change over time by setting parameters, its actual improvement in performance could be exponentially increased. Zhou Zhang 0008, Ying Liu 0024, Gang Ren 0003, Jun Bi |
LANMAN | 2 |
| 2015 | Building an IPv6 address generation and traceback system with NIDTGA in Address Driven Network
Ying Liu 0024, Gang Ren 0003, Shenglin Zhang, Lin He 0004, Yihao Jia |
Sci. China Inf. Sci. | 1 |
| 2015 | A bottleneck-free model for P4P
Ying Liu 0024, Shenglin Zhang |
Sci. China Inf. Sci. | 1 |
| 2015 | A Family of Stable Multipath Dual Congestion Control Algorithms
Ying Liu 0024, Ke Xu 0002, Meng Shen 0001 |
J. Comput. Sci. Technol. | 1 |
| 2014 | A measurement study on BGP AS path looping (BAPL) behaviorabstractAs a path vector protocol, Border Gateway Protocol (BGP) messages contain the entire Autonomous System (AS) path to each destination for breaking arbitrary long AS path loops. However, after observing the global routing data from RouteViews, we find that BGP AS path looping (BAPL) behavior does occur and in fact can lead to multi-AS forwarding loops in both IPv4 and IPv6. The number and ratio of BAPLs in IPv4 and IPv6 for 1456 days on a daily basis are analyzed. Moreover, the distribution of BAPL duration and loop length in IPv4 and IPv6 are also studied. Some possible explanations for BAPLs are discussed in this paper. Private AS number leaking has contributed to 1.76% of BAPLs in IPv4 and 0.00027% in IPv6, and at least 2.85% of BAPLs in IPv4 were attributed to faulty configurations and malicious attacks. Valid explanations, including multinational companies, preventing particular AS from accepting routes, can also lead to BAPLs. Shenglin Zhang, Ying Liu 0024, Dan Pei |
ICCCN | 2 |
| 2014 | CDRDN: Content Driven Routing in Datacenter NetworkabstractA major challenge in data center networks is to provide enough network capacity to meet the ever increasing demand of large-scale distributed computing. While existing proposals focus on adding more switches and links, which cost extra hardware and energy, we explore another dimension in which spare disk space at servers is used for caching data, so as to increase network throughput without additional hardware. Leveraging the Named Data Networking (NDN) architecture and unique characteristics of data centers, we design a novel Content Driven Routing in Datacenter Network (CDRDN) that enables universal caching in data center networks in an efficient and scalable way. First, rather than using switches for “on-path” caching like in native NDN, CDRDN uses the large storage space at servers to do “off-path” caching. Second, by taking advantage of data center's regular and hierarchical network topology, CDRDN switches are able to direct requests to nearby server caches even under high dynamics of caches. Third, given the vast amount of data in data centers, the full name-based routing table would be difficult to fit in switch's limited fast memory. CDRDN adopts a compound content and location routing to ensure packet delivery while benefiting from name-based routing as much as the switches can afford. CDRDN extends NDN's adaptive forwarding mechanism to deal with cache misses, link failures, and congestion without running routing protocols or cache exchange protocols. Our packet-level simulations show that CDRDN can almost double the network throughput compared with shortest-path routing under the same setting, and CDRDN can effectively deal with link failures using adaptive forwarding. Dan Li 0001, Ying Liu 0024 |
ICCCN | 3 |
| 2014 | TED: Inter-domain traffic engineering via deflectionabstractAs inter-domain routing on today's Internet does not and basically cannot consider traffic load when determining best traffic forwarding paths, it is not always optimal for a router to forward packets along its default path, especially when the router's default output port incurs a long queuing delay. In this paper, we design a new approach called TED in which border routers of autonomous systems (AS) adaptively deflect outbound traffic from a congested default path to an alternative path to significantly improve inter-domain traffic engineering (TE) and end-to-end throughput. With TED, every router only needs to examine the queue length of its own outgoing ports to orchestrate its deflection operation and ensure traffic forwarding is at line speed. It does not need to communicate or coordinate with other TED-capable routers or modify packet content, making TED incrementally deployable. Our evaluation shows that TED significantly increases the average throughput of traffic flows, and the improvement is comparable to directly upgrading router hardware and capacity. Finally, a prototype of TED on NetFPGA is also implemented. Jun Li 0001, Ying Liu 0024, Dan Li 0001 |
IWQoS | 3 |
| 2014 | CCOF: Congestion control on the fly for Inter-Domain routingabstractWe design a new routing system called CCOF in which border routers of autonomous systems (AS) adaptively redirect outbound traffic from a congested default path to an alternative path in favor of Inter Domain traffic engineering (TE). We show how this can be done simply, by examining outgoing port's queuing size, and efficiently, that maintains line speed forwarding. In addition, CCOF-router is completely compatible with legacy routers since it neither requires cooperation nor modifies the packet content. Our evaluation shows that CCOF significantly increases the average throughput of traffic flows, and has similar improvement as directly upgrading device capacity. Finally, prototype implementation is achieved on NetFPGA. Ying Liu 0024, Jun Li 0001 |
LANMAN | 2 |
| 2014 | Towards evolvable Internet architecture-design constraints and models analysis
Ke Xu 0002, Guangwu Hu, Yifeng Zhong, Ying Liu 0024, Ning Wang 0001 |
Sci. China Inf. Sci. | 6 |
| 2014 | Reliable Multicast in Data Center NetworksabstractMulticast benefits data center group communication in both saving network traffic and improving application throughput. Reliable packet delivery is required in data center multicast for data-intensive computations. However, existing reliable multicast solutions for the Internet are not suitable for the data center environment, especially with regard to keeping multicast throughput from degrading upon packet loss, which is norm instead of exception in data centers. We present RDCM, a novel reliable multicast protocol for data center network. The key idea of RDCM is to minimize the impact of packet loss on the multicast throughput, by leveraging the rich link resource in data centers. A multicast-tree-aware backup overlay is explicitly built on group members for peer-to-peer packet repair. The backup overlay is organized in such a way that it causes little individual repair burden, control overhead, as well as overall repair traffic. RDCM also realizes a window-based congestion control to adapt its sending rate to the traffic status in the network. Simulation results in typical data center networks show that RDCM can achieve higher application throughput and less traffic footprint than other representative reliable multicast protocols. We have implemented RDCM as a user-level library on Windows platform. The experiments on our test bed show that RDCM handles packet loss without obvious throughput degradation during high-speed data transmission, gracefully respond to link failure and receiver failure, and causes less than 10% CPU overhead to data center servers. Dan Li 0001, Mingwei Xu 0001, Ying Liu 0024, Yong Cui 0001, Guihai Chen |
IEEE Trans. Computers | 3 |
| 2013 | Utility function of TCP Reno under precise end-to-end drop probability
Ying Liu 0024 |
Sci. China Inf. Sci. | 1 |
| 2013 | Research achievements on the new generation Internet architecture and protocols
Ying Liu 0024, Zhou Zhang 0008, Ke Xu 0002 |
Sci. China Inf. Sci. | 1 |
| 2008 | Theoretical research progress in new-generation Internet architecture
Ying Liu 0024, Qian Wu 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2006 | Heterogeneous QoS Multicast and Its Improvement on Edge-Based Overlay Networks
Suogang Li, Ke Xu 0002, Ying Liu 0024 |
HPCC | 4 |
| 2006 | Building Trees to Support Comparable Multi-class Services in Edge Overlay MulticastabstractTraditional IP multicast in a network domain is likely to imply a huge burden of storage and forwarding for routers and it's hard to support quality of service. The recent proposed application layer multicast is more scalable but increases traffic load and end-to-end delay. In the paper, we make multicast supporting comparable multi-class services on the overlay network comprising only edge routers. Considering resource limitation on the router and multi-class services by the member, the problem to build minimum cost trees is NP-hard, so we design three feasible heuristic algorithms to solve it. Extensive simulations are conducted to evaluate the performance of the proposed heuristics and validate the effectiveness of reducing the total tree cost and iteration times under considered constraints. The proposal is expected to combine with DiffServ or MPLS VPN networks to fulfill multi-class QoS multicast. Suogang Li, Ke Xu 0002, Ying Liu 0024 |
ICCCN | 4 |
| 2006 | A modularized QoS multicasting approach on common homogeneous trees for heterogeneous members in DiffServabstractTraditional IP multicast suffers from forwarding state scalability problems as the number of concurrent active multicast groups increases. The problem is exacerbated when provisioning QoS since additional information of resource requirement from members must be kept at routers. In this paper, we firstly consider the comparability between various QoS levels and propose a modularized QoS multicasting approach in the DiffServ model. To achieve the multicast state scalability, the trees, called common trees, are decoupled from groups. The groups including the members requiring heterogeneous QoS are divided into subgroups and delivered through the common homogeneous QoS trees. We put forward a tree-choosing algorithm to find out appropriate trees for a group. Extensive simulations demonstrate that the approach is able to improve multicast forwarding state scalability as well as supporting different QoS requirement. The modularized approach can be considered a building block for QoS multicasting in the DiffServ architecture for its ease of deployment and compatibility with DiffServ Suogang Li, Ke Xu 0002, Ying Liu 0024 |
IPCCC | 4 |