Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Weikai Huang

dblp:288/7675 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0002-4059-7230ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 77% Language models and text generation · 23%
Network and information security
1 paper
Cryptographic primitives and cryptanalysis · 67% Biometric security · 33%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
cross-modal retrieval
1.012026
When Multimedia Meets Security: Privacy-Preserving Cross-Modal Retrieval for Large-Scale Data · IEEE Trans. Inf. Forensics Secur. 2026
Information retrieval › cross-modal retrieval
privacy-preserving cross-modal retrieval
1.012026
When Multimedia Meets Security: Privacy-Preserving Cross-Modal Retrieval for Large-Scale Data · IEEE Trans. Inf. Forensics Secur. 2026
Biometric security
cross-modal retrieval
1.012026
When Multimedia Meets Security: Privacy-Preserving Cross-Modal Retrieval for Large-Scale Data · IEEE Trans. Inf. Forensics Secur. 2026
Cryptographic primitives and cryptanalysis
searchable encryption
1.012026
When Multimedia Meets Security: Privacy-Preserving Cross-Modal Retrieval for Large-Scale Data · IEEE Trans. Inf. Forensics Secur. 2026
Cryptographic primitives and cryptanalysis › searchable encryption
searchable symmetric encryption
1.012026
When Multimedia Meets Security: Privacy-Preserving Cross-Modal Retrieval for Large-Scale Data · IEEE Trans. Inf. Forensics Secur. 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
0.812024
Task Me Anything · NeurIPS 2024
Natural language and speech › Language models and text generation › LLM agents
tool use
0.812024
m &m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks · ECCV (10) 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.212024
Task Me Anything · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

multi-key query · 2.0hamming inverted multi-index · 2.0adaptive security · 2.0tool-use agents · 0.8taxonomy construction · 0.8benchmark generation · 0.8benchmark evaluation · 0.8
YearPublicationVenuePosition
2026 When Multimedia Meets Security: Privacy-Preserving Cross-Modal Retrieval for Large-Scale Data
abstract
In recent years, Cross-Modal Retrieval (CMR), which can retrieve data across types based on query semantics, has become an attractive technology due to the widespread applications of multimedia data. Outsourcing multimedia data to a cloud server is a reliable way to improve the quality of CMR services, but it will also incur potential data privacy leakage issues. Existing schemes for privacy-preserving outsourced data search services are either inapplicable to CMR or limited by efficiency and scalability. To address the above issues, we investigate the problem of Searchable Symmetric Encryption (SSE) for CMR in this paper. Firstly, we formulate the definition of SSE for CMR (namely, SSECMR) and extend the SSE leakage functions to capture the leakage in SSECMR. Then, by constructing distance-computation-free Hamming inverted multi-index, we propose a practical SSECMRconstruction. Specifically, we transform Hamming distance-based range queries into multi-key queries, thereby avoiding computationally expensive comparison operations and enabling efficient Hamming distance queries over encrypted data. Our design supports secure CMR with sub-linear complexity in a single communication roundtrip. Through rigorous security analysis, we demonstrate that our construction can provide adaptive security. Empirical evaluations on real-world datasets demonstrate that SSECMRoutperforms the state-of-the-art scheme in both efficiency and accuracy, and is comparable to plaintext applications. Our code is available at https://github.com/SSE-CMR/SSECMR.
Weikai Huang, Xiangyu Wang 0010, Dan Zhu 0001, XinDi Ma, Jianfeng Ma 0001
IEEE Trans. Inf. Forensics Secur.1
2024 m &m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
Zixian Ma, Weikai Huang, Jieyu Zhang 0001, Tanmay Gupta, Ranjay Krishna
ECCV (10)2
2024 Task Me Anything
abstract
Benchmarks for large multimodal language models (MLMs) now serve to simultaneously assess the general capabilities of models instead of evaluating for a specific capability. As a result, when a developer wants to identify which models to use for their application, they are overwhelmed by the number of benchmarks and remain uncertain about which benchmark's results are most reflective of their specific use case. This paper introduces Task-Me-Anything, a benchmark generation engine which produces a benchmark tailored to a user's needs. Task-Me-Anything maintains an extendable taxonomy of visual assets and can programmatically generate a vast number of task instances. Additionally, it algorithmically addresses user queries regarding MLM performance efficiently within a computational budget. It contains 113K images, 10K videos, 2K 3D object assets, over 365 object categories, 655 attributes, and 335 relationships. It can generate 500M image/video question-answering pairs, which focus on evaluating MLM perceptual capabilities. Task-Me-Anything reveals critical insights: open-source MLMs excel in object and attribute recognition but lack spatial and temporal understanding; each model exhibits unique strengths and weaknesses; larger models generally perform better, though exceptions exist; and GPT4O demonstrates challenges in recognizing rotating/moving objects and distinguishing colors.
Jieyu Zhang 0001, Weikai Huang, Zixian Ma, Oscar Michel, Dong He 0002, Tanmay Gupta, Wei-Chiu Ma, Ali Farhadi, Aniruddha Kembhavi, Ranjay Krishna
NeurIPS2