VLDB 2026 Research / reviewers in the wild / expert
Nan Hua
dblp:28/5824
· DBLP profile ↗
22ranked-venue papers
9as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 8 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Graders Should Cheat: Privileged Information Enables Expert-Level Automated EvaluationsabstractAuto-evaluating language models (LMs), i.e., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and reduce the cost associated with it.But this presents a paradox: how can we trust the grader LM, which is presumably weaker than the candidate LM, to assess problems that are beyond the frontier of the capabilities of either model or both?For instance, today's LMs struggle on graduate-level physics and Olympiad-level math, making them unreliable graders in these domains.We show that providing privileged information -such as ground-truth solutions or problem-specific guidelines -improves automated evaluations on such frontier problems.This approach offers two key advantages.First, it expands the range of problems where LMs graders apply.Specifically, weaker models can now rate the predictions of stronger models.Second, privileged information can be used to devise easier variations of challenging problems which improves the separability of different LMs on tasks where their performance is generally low.With this approach, general-purpose LM graders match the state of the art performance on RewardBench, surpassing almost all the specially-tuned models.LM graders also outperform individual human raters on Vibe-Eval, and approach human expert graders on Olympiad-level math problems. Jin Peng Zhou, Sébastien M. R. Arnold, Nan Ding 0002, Kilian Q. Weinberger, Nan Hua, Fei Sha |
EMNLP | 5 |
| 2023 | FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information ExtractionabstractChen-Yu Lee, Chun-Liang Li, Hao Zhang, Timothy Dozat, Vincent Perot, Guolong Su, Xiang Zhang, Kihyuk Sohn, Nikolay Glushnev, Renshen Wang, Joshua Ainslie, Shangbang Long, Siyang Qin, Yasuhisa Fujii, Nan Hua, Tomas Pfister. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Kihyuk Sohn, Nikolay Glushnev, Renshen Wang, Joshua Ainslie, Shangbang Long, Siyang Qin, Yasuhisa Fujii, Nan Hua, Tomas Pfister |
ACL (1) | 15 |
| 2023 | Notice the Imposter! A Study on User Tag Spoofing Attack in Mobile Apps
Shuai Li 0006, Zhemin Yang, Guangliang Yang 0001, Hange Zhang, Nan Hua, Yurui Huang, Min Yang 0002 |
USENIX Security Symposium | 5 |
| 2022 | FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information ExtractionabstractChen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, Tomas Pfister. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, Tomas Pfister |
ACL (1) | 6 |
| 2022 | Collect Responsibly But Deliver Arbitrarily?: A Study on Cross-User Privacy Leakage in Mobile AppsabstractRecent years have witnessed the interesting trend that modern mobile apps perform more and more likely as user-to-user platforms, where app users can be freely and conveniently connected. Upon these platforms, rich and diverse data is often delivered across users, which brings users great conveniences and plentiful services, but also introduces privacy security concerns. While prior work has primarily studied illegitimate personal data collection problems in mobile apps, few paid little attention to the security of this emerging user-to-user platform feature, thus providing a rather limited understanding of the privacy risks in this aspect. Shuai Li 0006, Zhemin Yang, Nan Hua, Peng Liu 0005, Xiaohan Zhang 0001, Guangliang Yang 0001, Min Yang 0002 |
CCS | 3 |
| 2022 | Protoformer: Embedding Prototypes for Transformers
Ashkan Farhangi, Ning Sui, Nan Hua, Haiyan Bai, Arthur Huang, Zhishan Guo |
PAKDD (1) | 3 |
| 2019 | Provisioning Short-Term Traffic Fluctuations in Elastic Optical NetworksabstractTransient traffic spikes are becoming a crucial challenge for network operators from both user-experience and network-maintenance perspectives. Different from long-term traffic growth, the bursty nature of short-term traffic fluctuations makes it difficult to be provisioned effectively. Luckily, next-generation elastic optical networks (EONs) provide an economical way to deal with such short-term traffic fluctuations. In this paper, we go beyond conventional network reconfiguration approaches by proposing the novel lightpath-splitting scheme in EONs. In lightpath splitting, we introduce the concept of SplitPoints to describe how lightpath splitting is performed. Lightpaths traversing multiple nodes in the optical layer can be split into shorter ones by SplitPoints to serve more traffic demands by raising signal modulation levels of lightpaths accordingly. We formulate the problem into a mathematical optimization model and linearize it into an integer linear program (ILP). We solve the optimization model on a small network instance and design scalable heuristic algorithms based on greedy and simulated annealing approaches. Numerical results show the tradeoff between throughput gain and negative impacts like traffic interruptions. Especially, by selecting SplitPoints wisely, operators can achieve almost twice as much throughput as conventional schemes without lightpath splitting. Zhizhen Zhong, Nan Hua, Massimo Tornatore, Jialong Li 0006, Yanhe Li, Xiaoping Zheng, Biswanath Mukherjee |
IEEE/ACM Trans. Netw. | 2 |
| 2018 | Andromeda: Performance, Isolation, and Velocity at Scale in Cloud Network Virtualization
Michael Dalton, David Schultz, Jacob Adriaens, Ahsan Arefin, Anshuman Gupta, Brian Fahs, Dima Rubinstein, Enrique Cauich Zermeno, Erik Rubow, James Alexander Docauer, Jesse Alpert, Jing Ai, Jon Olson, Kevin DeCabooter, Marc de Kruijf, Nan Hua, Nathan Lewis, Nikhil Kasinadhuni, Riccardo Crepaldi, Srinivas Krishnan, Subbaiah Venkata, Yossi Richter, Uday Naik, Amin Vahdat |
NSDI | 16 |
| 2016 | On QoS-Assured Degraded Provisioning in Service-Differentiated Multi-Layer Elastic Optical NetworksabstractDegraded provisioning provides an effective solution to flexibly allocate resources in various dimensions to reduce blocking for differentiated demands when network congestion occurs. In this work, we investigate the novel problem of online degraded provisioning in service-differentiated multi-layer networks with optical elasticity. Quality of Service (QoS) is assured by service-holding-time prolongation and immediate access as soon as the service arrives without set-up delay. We decompose the problem into degraded routing and degraded resource allocation stages, and design polynomial-time algorithms with the enhanced multi-layer architecture to exploit network flexibility in temporal and spectral dimensions. Numerical results verify that we can achieve significant blocking reduction, especially for requests with higher priorities. They also indicate that degradation in optical layer can increase the network capacity, while degradation in electric layer provides flexible time-bandwidth exchange. Zhizhen Zhong, Jipu Li, Nan Hua, Gustavo B. Figueiredo, Yanhe Li, Xiaoping Zheng, Biswanath Mukherjee |
GLOBECOM | 3 |
| 2012 | A simpler and better design of error estimating codingabstractWe study error estimating codes with the goal of establishing better bounds for the theoretical and empirical overhead of such schemes. We explore the idea of using sketch data structures for this problem, and show that the tug-of-war sketch gives an asymptotically optimal solution. The optimality of our algorithms are proved using communication complexity lower bound techniques. We then propose a novel enhancement of the tug-of-war sketch that greatly reduces the communication overhead for realistic error rates. Our theoretical analysis and assertions are supported by extensive experimental evaluation. Nan Hua, Ashwin Lall, Baochun Li, Jun (Jim) Xu |
INFOCOM | 1 |
| 2012 | Towards optimal error-estimating codes through the lens of Fisher information analysisabstractError estimating coding (EEC) has recently been established as an important tool to estimate bit error rates in the transmission of packets over wireless links, with a number of potential applications in wireless networks. In this paper, we present an in-depth study of error estimating codes through the lens of Fisher information analysis and find that the original EEC estimator fails to exploit the information contained in its code to the fullest extent. Motivated by this discovery, we design a new estimator for the original EEC algorithm, which significantly improves the estimation accuracy, and is empirically very close to the Cramer-Rao bound. Following this path, we generalize the EEC algorithm to a new family of algorithms called gEEC generalized EEC. These algorithms can be tuned to hold 25-35% more information with the same overhead, and hence deliver even better estimation accuracy---close to optimal, as evidenced by the Cramer-Rao bound. Our theoretical analysis and assertions are supported by extensive experimental evaluation. Nan Hua, Ashwin Lall, Baochun Li, Jun (Jim) Xu |
SIGMETRICS | 1 |
| 2011 | Non-crypto Hardware Hash Functions for High Performance Networking ASICsabstractHash functions are vital in networking. Hash-based algorithms are increasingly deployed in mission-critical, high speed network devices. These devices will need small, quick, hardware hash functions to keep up with Internet growth. There are many hardware hash functions used in this situation, foremost among them CRC-32. We develop parametrized methods for evaluating hash function output quality so as to better compare similar hash functions. We use these methods to explore the quality of candidate hash functions, including CRC-32, H3(with fixed seed), MD5 and others. We also propose optimized building blocks for hardware hash functions based on SP-networks. Given a size budget of 4K gates and only 1 cycle to compute the result, we demonstrate a 128 bit input, 64 bit output hash function built using this framework that ranks highly in our tests. Nan Hua, Eric Norige, Sailesh Kumar, Bill Lynch |
ANCS | 1 |
| 2011 | Towards a Universal Sketch for Origin-Destination Network Measurements
Haiquan (Chuck) Zhao, Nan Hua, Ashwin Lall, Ping Li 0001, Jia Wang 0001, Jun (Jim) Xu |
NPC | 2 |
| 2011 | BRICK: a novel exact active statistics counter architectureabstractIn this paper, we present an exact active statistics counter architecture called Bucketized Rank Indexed Counters (BRICK) that can efficiently store per-flow variable-width statistics counters entirely in SRAM while supporting both fast updates and lookups (e.g., 40-Gb/s line rates). BRICK exploits statistical multiplexing by randomly bundling counters into small fixed-size buckets and supports dynamic sizing of counters by employing an innovative indexing scheme called rank indexing. Experiments with Internet traces show that our solution can indeed maintain large arrays of exact active statistics counters with moderate amounts of SRAM. Nan Hua, Jun (Jim) Xu, Bill Lin 0001, Haiquan (Chuck) Zhao |
IEEE/ACM Trans. Netw. | 1 |
| 2010 | Automatic text categorization based on content analysis with cognitive situation models
Yi Guo 0009, Zhiqing Shao, Nan Hua |
Inf. Sci. | 3 |
| 2010 | A cognitive interactionist sentence parser with simple recurrent networks
Yi Guo 0009, Zhiqing Shao, Nan Hua |
Inf. Sci. | 3 |
| 2009 | Variable-Stride Multi-Pattern Matching For Scalable Deep Packet InspectionabstractAbstract—Accelerating multi-pattern matching is a critical is-sue in building high-performance deep packet inspection systems. Achieving high-throughputs while reducing both memory-usage and memory-bandwidth needs is inherently difficult. In this paper, we propose a pattern (string) matching algorithm that achieves high throughput while limiting both memory-usage and memory-bandwidth. We achieve this by moving away from a byte-oriented processing of patterns to a block-oriented scheme. However, different from previous block-oriented approaches, our scheme uses variable-stride blocks. These blocks can be uniquely identified in both the pattern and the input stream, hence avoid-ing the multiplied memory costs which is intrinsic in previous approaches. We present the algorithm, tradeoffs, optimizations, and implementation details. Performance evaluation is done using the Snort and ClamAV pattern sets. Using our algorithm, the throughput of a single search engine can easily have a many-fold increase at a small storage cost, typically less than three bytes per pattern character. I. Nan Hua, Haoyu Song 0001, T. V. Lakshman |
INFOCOM | 1 |
| 2008 | BRICK: a novel exact active statistics counter architectureabstractIn this paper, we present an exact active statistics counter architecture called BRICK (Bucketized Rank Indexed Counters) that can efficiently store per-flow variable-width statistics counters entirely in SRAM while supporting both fast updates and lookups (e.g., 40 Gb/s line rates). BRICK exploits statistical multiplexing by randomly bundling counters into small fixed-size buckets and supports dynamic sizing of counters by employing an innovative indexing scheme called rank-indexing. Experiments with Internet traces show that our solution can indeed maintain large arrays of exact active statistics counters with moderate amounts of SRAM. Nan Hua, Bill Lin 0001, Jun (Jim) Xu, Haiquan (Chuck) Zhao |
ANCS | 1 |
| 2008 | Packet doppler: network monitoring using packet shift detectionabstractDue to recent large-scale deployments of delay and loss-sensitive applications, there are increasingly stringent demands on the monitoring of service level agreement metrics. Although many end-to-end monitoring methods have been proposed, they are mainly based on active probing and thus inject measurement traffic into the network. In this paper, we propose a new scheme for monitoring service level agreement metrics, in particular, delay distribution. Our scheme is passive and therefore will not cause perturbation to real traffic. Using realistic delay and traffic demands, we show that our scheme achieves high accuracy and can detect burst events that will be missed by probing based methods. Tongqing Qiu, Jian Ni, Hao Wang 0010, Nan Hua, Yang Richard Yang, Jun (Jim) Xu |
CoNEXT | 4 |
| 2008 | Rank-indexed hashing: A compact construction of Bloom filters and variantsabstractBloom filter and its variants have found widespread use in many networking applications. For these applications, minimizing storage cost is paramount as these filters often need to be implemented using scarce and costly (on-chip) SRAM. Besides supporting membership queries, Bloom filters have been generalized to support deletions and the encoding of information. Although a standard Bloom filter construction has proven to be extremely space-efficient, it is unnecessarily costly when generalized. Alternative constructions based on storing fingerprints in hash tables have been proposed that offer the same functionality as some Bloom filter variants, but using less space. In this paper, we propose a new fingerprint hash table construction called Rank-Indexed Hashing that can achieve very compact representations. A rank-indexed hashing construction that offers the same functionality as a counting Bloom filter can be achieved with a factor of three or more in space savings even for a false positive probability of just 1%. Even for a basic Bloom filter function that only supports membership queries, a rank-indexed hashing construction requires less space for a false positive probability as high as 0.1%, which is significant since a standard Bloom filter construction is widely regarded as extremely space-efficient for approximate membership problems. Nan Hua, Haiquan (Chuck) Zhao, Bill Lin 0001, Jun (Jim) Xu |
ICNP | 1 |
| 2007 | Simple and Fair Scheduling Algorithm for Combined Input-Crosspoint-Queued SwitchabstractWe propose a fair and simple high-performance scheduling algorithm for combined input-crosspoint-queued switches, which is called tracking fair quota allocation (TFQA). Our algorithm is based on low-cost round-Robin scheme, which prioritizes the ports lagging behind our fair quota allocation scheme. Simulation shows that our algorithm could maintain over 99% throughput and achieve relatively low mean delay under almost all typical test traffic patterns, outperforming all known algorithms with the same implementation complexity, especially under heavy load scenarios. Moreover, our algorithm could provide max-min fairness under inadmissible traffic, better than many other typical algorithms proposed before. Nan Hua, Depeng Jin, Lieguang Zeng, Gang Feng 0004 |
ICC | 1 |
| 2006 | A Practical Switch-Memory-Switch Architecture Emulating PIFO OQabstractEmulating Output Queued (OQ) Switch with sustainable implementation cost and low fixed delay is always preferable in designing high performance routers. The Switch-Memory-Switch (SMS) router, also called Distributed Shared Memory (DSM) Switch, provides a possible way towards practically emulating OQ in backbone switches. However, the architectures and algorithms for SMS switches ever proposed are either unpractical or only supporting First-Come-First-Serve (FCFS) scheduling policy, which cannot support QoS and is unfair for light traffic flow. Our improved SMS architecture and algorithm aim at emulating Push-In-First-Out (PIFO) OQ. We employ a randomly-dispatching first stage and resolve memory access conflictions on the second stage of the switch through a probabilistic matching method, at the cost of fixed delay and sufficiently low cell loss probability (PCLP). The relative fixed delay of our algorithms for an NXN switch is composed of two parts: N and (-3/2log2PCLP), which result from the pipelined scheduling process and probabilistic method, respectively. Moreover, both the total memory and fabric bandwidth of our architecture implemented on crossbar could be lowered to only 2NR, where R is line rate, counting read and write separately. Nan Hua, Yang Xu 0010, Depeng Jin, Lieguang Zeng |
GLOBECOM | 1 |