VLDB 2026 Research / reviewers in the wild / expert
Huaming Rao
dblp:137/3711
· DBLP profile ↗
8ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorArtificial intelligence and machine learning · 1Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
2 papers |
Cryptographic protocols and secure computation · 65% Privacy and data protection · 35% | |
| Databases, data mining, and information retrieval
3 papers |
Machine learning and data management · 52% Data integration and cleaning · 39% Web and social media mining · 8% | |
| Human-computer interaction and pervasive computing
2 papers |
Collaborative and social computing · 52% Design research and methods · 48% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cryptographic protocols and secure computation › private set intersection
circuit-PSI |
1.3 | 2 | 2026 | Practical Anonymous Two-Party Gradient Boosting Decision Tree · SP 2026 Privacy-Preserving Screening for Record Linkage · ICDE 2025 |
Cryptographic protocols and secure computation
private set intersection |
1.3 | 2 | 2026 | Practical Anonymous Two-Party Gradient Boosting Decision Tree · SP 2026 Privacy-Preserving Screening for Record Linkage · ICDE 2025 |
Privacy and data protection
privacy-preserving machine learning |
1.0 | 1 | 2026 | Practical Anonymous Two-Party Gradient Boosting Decision Tree · SP 2026 |
Cryptographic protocols and secure computation
secure multiparty computation |
1.0 | 1 | 2026 | Practical Anonymous Two-Party Gradient Boosting Decision Tree · SP 2026 |
Data integration and cleaning
data preprocessing |
0.9 | 1 | 2025 | DataLab: A Unified Platform for LLM-Powered Business Intelligence · ICDE 2025 |
Visualization and visual analytics › visualization generation
automated visualization generation |
0.9 | 1 | 2025 | DataLab: A Unified Platform for LLM-Powered Business Intelligence · ICDE 2025 |
Privacy and data protection
privacy-preserving record linkage |
0.9 | 1 | 2025 | Privacy-Preserving Screening for Record Linkage · ICDE 2025 |
Collaborative and social computing › crowdsourcing
crowd feedback |
0.2 | 1 | 2015 | A Classroom Study of Using Crowd Feedback in the Iterative Design Process · CSCW 2015 |
Design research and methods › design process
design feedback |
0.2 | 1 | 2015 | A Classroom Study of Using Crowd Feedback in the Iterative Design Process · CSCW 2015 |
Design research and methods › design process
iterative design |
0.2 | 1 | 2015 | A Classroom Study of Using Crowd Feedback in the Iterative Design Process · CSCW 2015 |
Collaborative and social computing
crowdfunding |
0.2 | 1 | 2014 | Show me the money!: an analysis of project updates during crowdfunding campaigns · CHI 2014 |
Collaborative and social computing
online communities |
0.1 | 1 | 2014 | Show me the money!: an analysis of project updates during crowdfunding campaigns · CHI 2014 |
Methods — techniques the papers use, named apart from their topics
oblivious programmable pseudorandom function · 2.0homomorphic encryption · 2.0large language model · 1.7inter-agent communication · 1.7domain knowledge incorporation · 1.7agent framework · 1.7secure permutation · 0.9oblivious attribute alignment · 0.9circuit-PSI · 0.9semantic analysis · 0.4corpus analysis · 0.4expert evaluation · 0.2classroom study · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Practical Anonymous Two-Party Gradient Boosting Decision TreeabstractStructured data is well handled by gradient-boosted decision trees (GBDT), which are usually trained on vertically partitioned features across mutually distrustful parties. High speed and interpretability make GBDTs popular in finance and healthcare, where neural networks may fall short. Enabling secure computation for GBDTs poses unique challenges, requiring secure record alignment for comparison. Relying on private set intersection (PSI) is a de facto approach. Mistaking PSI for a safety measure actually exposes which record identifiers (IDs) are shared between the datasets. Although circuit-PSI could help, it is costly for generic uses. New ideas are needed to efficiently train in a "dark forest". Aiming to hide the IDs, we initiate the study of anonymous GBDT training on split data held by two parties. Dual circuit-PSI in our design lets the parties alternate as receiver to run pick-then-sum over local features. Via oblivious programmable pseudorandom functions, we propagate circuit-PSI outputs as shared state across runs. Avoiding universal alignment, we resolve the neglected dilemma that ID hiding incurs a cost that scales with domain size. Next, we halve the cost of ciphertext packing used to convert single-instruction multiple-data homomorphic encryption from (ring) learning with errors in prior secure GBDT (Usenix Security' 23) and related secure machine-learning computations. Comparative experiments show our protocol remains competitive with leaky approaches in efficiency. Enabling ID-hiding aggregation, our techniques can extend to other vertically partitioned analytics. Minxin Du, Sherman S. M. Chow, Huangxun Chen, Huaming Rao, Danqing Huang, Peng Chen 0021 |
SP | 6 |
| 2025 | Privacy-Preserving Screening for Record LinkageabstractIn an era dominated by big data and machine learning, establishing valuable data collaboration has never been more critical. However, such collaborations must operate under regulatory and legal constraints. Two-party Privacy-Preserving Record Linkage (PPRL) emerges to assess the potential collaboration value and also ensure the privacy and security of the involved data. Nevertheless, the substantial computational and communication overheads associated with PPRL hinder its practical adoption in data markets with numerous potential collaborators. Therefore, we present the Screening-then-Linkage framework, which incorporates a lightweight Screening phase prior to the resource-intensive PPRL phase, i.e., PPRS, to mitigate the scalability issue of PPRL. We propose a circuit-PSI-based system, named Appraisal to realize a secure, effective, and efficient PPRS. To reconcile the approximate matching and/or schema-aware setting required in PPRS with the limitations of the circuit-PSI supporting only symmetric functions, we propose a more communication-efficient secure permutation, i.e., Oblivious Attribute/Feature Alignment protocol tailored for PPRS. This protocol supports a broader range of comparison functions and significantly improves efficiency, i.e., reducing communication costs by a factor of 14 compared to the conventional protocol. Our rigorous analysis and comprehensive empirical evaluations demonstrate the security, effectiveness, and efficiency of Appraisal. Appraisal can accommodate up to 850x more records than the SOTA PPRS system, SFour, within the same constraints. Moreover, it is 165x faster than SOTA PPRL, indicating the Screening-then-Linkage framework substantially decreases the computation time required to identify the most valuable collaborators from a large pool of candidates. Huangxun Chen, Yongjun Zhao 0001, Huaming Rao, Danqing Huang |
ICDE | 5 |
| 2025 | DataLab: A Unified Platform for LLM-Powered Business IntelligenceabstractBusiness intelligence (BI) transforms large volumes of data within modern organizations into actionable insights for informed decision-making. Recently, large language model (LLM)-based agents have streamlined the BI workflow by automatically performing task planning, reasoning, and actions in executable environments based on natural language (NL) queries. However, existing approaches primarily focus on individual BI tasks such as NL2SQL and NL2VIS. The fragmentation of tasks across different data roles and tools lead to inefficiencies and potential errors due to the iterative and collaborative nature of BI. In this paper, we introduce DataLab, a unified BI platform that integrates a one-stop LLM-based agent framework with an augmented computational notebook interface. DataLab supports various BI tasks for different data roles in data preparation, analysis, and visualization by seamlessly combining LLM assistance with user customization within a single environment. To achieve this unification, we design a domain knowledge incorporation module tailored for enterprise-specific BI tasks, an inter-agent communication mechanism to facilitate information sharing across the BI workflow, and a cell-based context management strategy to enhance context utilization efficiency in BI notebooks. Extensive experiments demonstrate that DataLab achieves state-of-the-art performance on various BI tasks across popular research benchmarks. Moreover, DataLab maintains high effectiveness and efficiency on real-world datasets from Tencent, achieving up to a 58.58% increase in accuracy and a 61.65 % reduction in token cost on enterprise-specific BI tasks. Luoxuan Weng, Yinghao Tang, Yingchaojie Feng, Zhuo Chang, Ruiqin Chen, Haozhe Feng, Chen Hou, Danqing Huang, Yang Li 0106, Huaming Rao, Canshi Wei, Xiuqi Huang, Minfeng Zhu 0001, Yuxin Ma 0001, Bin Cui 0001, Peng Chen 0021, Wei Chen 0001 |
ICDE | 10 |
| 2020 | CAN-GAN: Conditioned-attention normalized GAN for face age synthesis
Chenglong Shi, Jiachao Zhang, Yazhou Yao, Yunlian Sun, Huaming Rao, Xiangbo Shu |
Pattern Recognit. Lett. | 5 |
| 2016 | Leveraging Human Computations to Improve Schematization of Spatial Relations from ImageryabstractThe process of generating schematic maps of salient objects from a set of pictures of an indoor environment is challenging. It has been an active area of research as it is crucial to a wide range of context- and location-aware services, as well as for general scene understanding. Although many automated systems have been developed to solve the problem, most of them either require predefining labels or expensive equipment, such as RGBD sensors or lasers, to scan the environment. In this article, we introduce a prototype system to show how human computations can be utilized to generate schematic maps from a set of pictures, without making strong assumptions or demanding extra devices. The system requires humans (crowd workers from Amazon Mechanical Turks) to do simple spatial mapping tasks in various conditions, and their data are aggregated by filtering and clustering techniques that allow salient cues to be identified in the pictures and their spatial relations to be inferred and projected on a two-dimensional map. In particular, we tested and demonstrated the effectiveness of two methods that improved the quality of the generated schematic map: (1) We encouraged humans to adopt an allocentric representations of salient objects by guiding them to perform mental rotations of these objects and (2) we sensitized human perception by guided arrows superimposed on the imagery to improve the accuracy of depth and width estimation. We demonstrated the feasibility of our system by evaluating the results of schematic maps generated from indoor pictures taken from an office building. By calculating Riemannian shape distances between the generated maps to the ground truth, we found that the generated schematic maps captured the spatial relations well. Our results showed that the combination of human computations and machine clustering could lead to more-accurate schematized maps from imagery. We also discuss how our approach may have important insights on methods that leverage human computations in other areas. Huaming Rao, Shih-Wen Huang, Wai-Tat Fu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2015 | A Classroom Study of Using Crowd Feedback in the Iterative Design ProcessabstractCrowd feedback systems offer designers an emerging approach for improving their designs, but there is little empirical evidence of the benefit of these systems. This paper reports the results of a study of using a crowd feedback system to iterate on visual designs. Users in an introductory visual design course created initial designs satisfying a design brief and received crowd feedback on the designs. Users revised the designs and the system was used to generate feedback again. This format enabled us to detect the changes between the initial and revised designs and how the feedback related to those changes. Further, we analyzed the value of crowd feedback by comparing it with expert evaluation and feedback generated via free-form prompts. Results showed that the crowd feedback system prompted deep and cosmetic changes and led to improved designs, the crowd recognized the design improvements, and structured workflows generated more interpretative, diverse and critical feedback than free-form prompts. Anbang Xu, Huaming Rao, Steven Dow, Brian P. Bailey |
CSCW | 2 |
| 2014 | Show me the money!: an analysis of project updates during crowdfunding campaignsabstractHundreds of thousands of crowdfunding campaigns have been launched, but more than half of them have failed. To better understand the factors affecting campaign outcomes, this paper targets the content and usage patterns of project updates -- communications intended to keep potential funders aware of a campaign's progress. We analyzed the content and usage patterns of a large corpus of project updates on Kickstarter, one of the largest crowdfunding platforms. Using semantic analysis techniques, we derived a taxonomy of the types of project updates created during campaigns, and found discrepancies between the design intent of a project update and the various uses in practice (e.g. social promotion). The analysis also showed that specific uses of updates had stronger associations with campaign success than the project's description. Design implications were formulated from the results to help designers better support various uses of updates in crowdfunding campaigns. Anbang Xu, Huaming Rao, Wai-Tat Fu, Shih-Wen Huang, Brian P. Bailey |
CHI | 3 |
| 2013 | What Will Others Choose? How a Majority Vote Reward Scheme Can Improve Human Computation in a Spatial Location Identification TaskabstractWe created a spatial location identification task (SpLIT) in which workers recruited from Amazon Mechanical Turk were presented with a camera view of a location, and were asked to identify the location on a two-dimensional map. In cases where these cues were ambiguous or did not provide enough information to pinpoint the exact location, workers had to make a best guess. We tested the effects of two reward schemes. In the “ground truth” scheme, workers were rewarded if their answers were close enough to the correct locations. In the “majority vote” scheme, workers were told that they would be rewarded if their answers were similar to the majority of other workers. Results showed that the majority vote reward scheme led to consistently more accurate answers. Cluster analysis further showed that the majority vote reward scheme led to answers with higher reliability (a higher percentage of answers in the correct clusters) and precision (a smaller average distance to the cluster centers). Possible reasons for why the majority voting reward scheme was better were discussed. Huaming Rao, Shih-Wen Huang, Wai-Tat Fu |
HCOMP | 1 |