VLDB 2026 Research / reviewers in the wild / expert
Binxing Jiao
dblp:78/2418
· DBLP profile ↗
12ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-4710-0095ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
7 papers |
Information retrieval · 100% | |
| Artificial intelligence
7 papers |
Representation and self-supervised learning · 71% Language models and text generation · 11% Speech recognition and synthesis · 11% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 24 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval models |
2.4 | 4 | 2023 | LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval · WWW 2023 LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval · ICLR 2023 Effective and Efficient Query-aware Snippet Extraction for Web Search · EMNLP 2022 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
1.8 | 3 | 2023 | LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval · WWW 2023 LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval · ICLR 2023 xMoCo: Cross Momentum Contrastive Learning for Open-Domain Question Answering · ACL/IJCNLP (1) 2021 |
Machine learning › Representation and self-supervised learning
contrastive learning |
1.2 | 2 | 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026 LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval · ICLR 2023 |
Machine learning › Representation and self-supervised learning › text embedding
text representation learning |
1.0 | 1 | 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026 |
Machine learning › Representation and self-supervised learning › text embedding › sentence embedding
unsupervised sentence embeddings |
1.0 | 1 | 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026 |
Program synthesis and code generation
code agent |
1.0 | 1 | 2026 | GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging · AAAI 2026 |
Machine learning › Representation and self-supervised learning
pre-training |
0.9 | 2 | 2023 | LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval · ICLR 2023 SimLM: Pre-training with Representation Bottleneck for Dense Passage Retrieval · ACL (1) 2023 |
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining |
0.8 | 1 | 2024 | VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation Learning · IEEE Trans. Multim. 2024 |
Natural language and speech › Speech recognition and synthesis
speech representation learning |
0.8 | 1 | 2024 | VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation Learning · IEEE Trans. Multim. 2024 |
Information retrieval › document retrieval › passage retrieval
dense passage retrieval |
0.7 | 1 | 2023 | SimLM: Pre-training with Representation Bottleneck for Dense Passage Retrieval · ACL (1) 2023 |
Information retrieval › retrieval models
knowledge distillation for retrieval |
0.7 | 1 | 2023 | LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval · WWW 2023 |
Security and privacy of machine learning
model intellectual property protection |
0.7 | 1 | 2023 | Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark · ACL (1) 2023 |
Security and privacy of machine learning › model intellectual property protection
model watermarking |
0.7 | 1 | 2023 | Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark · ACL (1) 2023 |
Information retrieval
search engines |
0.6 | 2 | 2022 | Effective and Efficient Query-aware Snippet Extraction for Web Search · EMNLP 2022 Visually Summarizing Web Pages Through Internal and External Images · IEEE Trans. Multim. 2012 |
Information retrieval › text summarization
query-biased snippet generation |
0.6 | 1 | 2022 | Effective and Efficient Query-aware Snippet Extraction for Web Search · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.5 | 1 | 2021 | xMoCo: Cross Momentum Contrastive Learning for Open-Domain Question Answering · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation › large language model
large language model adaptation |
0.3 | 1 | 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026 |
Natural language and speech › Language models and text generation
LLM agents |
0.3 | 1 | 2026 | GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging · AAAI 2026 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked prediction |
0.2 | 1 | 2024 | VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation Learning · IEEE Trans. Multim. 2024 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2023 | Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark · ACL (1) 2023 |
Information retrieval
image retrieval |
0.1 | 1 | 2012 | Visually Summarizing Web Pages Through Internal and External Images · IEEE Trans. Multim. 2012 |
Information retrieval › search interfaces
search result presentation |
0.1 | 1 | 2012 | Visually Summarizing Web Pages Through Internal and External Images · IEEE Trans. Multim. 2012 |
Information retrieval › interactive information retrieval › search tasks
re-finding |
0.1 | 1 | 2010 | Visual summarization of web pages · SIGIR 2010 |
Multimedia analysis and retrieval
image retrieval |
0.0 | 1 | 2010 | Visual summarization of web pages · SIGIR 2010 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 3.0pre-training · 2.6benchmark evaluation · 2.0embedding-based watermark · 1.3backdoor watermarking · 1.3masked next-token prediction · 1.0context compression · 1.0unified tokenizer · 0.8masked prediction · 0.8cross-modal representation learning · 0.8representation bottleneck · 0.7rank-consistent regularization · 0.7knowledge distillation · 0.7two-stage coarse-to-fine inference · 0.6neural sentence representation · 0.6momentum contrastive learning · 0.5cross momentum · 0.5user study · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository LeveragingabstractBeyond scratch coding, exploiting large-scale code repositories (e.g., GitHub) for practical tasks is vital in real-world software development, yet current benchmarks rarely evaluate code agents in such authentic, workflow-driven scenarios. To bridge this gap, we introduce GitTaskBench, a benchmark designed to systematically assess this capability via 54 realistic tasks across 7 modalities and 7 domains. Each task pairs a relevant repository with an automated, human-curated evaluation harness specifying practical success criteria. Beyond measuring execution and task success, we also propose the alpha-value metric to quantify the economic benefit of agent performance, which integrates task success rates, token cost, and average developer salaries. Experiments across three state-of-the-art agent frameworks with multiple advanced LLMs show that leveraging code repositories for complex task solving remains challenging: even the best-performing system, OpenHands+Claude 3.7, solves only 48.15% of tasks. Error analysis attributes over half of failures to seemingly mundane yet critical steps like environment setup and dependency resolution, highlighting the need for more robust workflow management and increased timeout preparedness. By releasing GitTaskBench, we aim to drive progress and attention toward repository-aware code reasoning, execution, and deployment---moving agents closer to solving complex, end-to-end real-world tasks. Ziyi Ni, Huacan Wang, Shuo Lu, Wang You, Zhenheng Tang, Sen Hu 0005, Bo Li 0117, Binxing Jiao, Daxin Jiang, Yuntao Du 0001 |
AAAI | 11 |
| 2026 | Learning to Compress: Unlocking the Potential of Large Language Models for Text RepresentationabstractText representation plays a critical role in tasks like clustering, retrieval, and other downstream applications. With the emergence of large language models (LLMs), there is increasing interest in harnessing their capabilities for this purpose. However, most of the LLMs are inherently causal and optimized for next-token prediction, making them suboptimal for producing holistic representations. To address this, recent studies introduced pretext tasks to adapt LLMs for text representation. Most of these tasks, however, rely on token-level prediction objectives, such as the masked next-token prediction (MNTP) used in LLM2Vec. In this work, we explore the untapped potential of context compression as a pretext task for unsupervised adaptation of LLMs. During compression pre-training, the model learns to generate compact memory tokens, which substitute the whole context for downstream sequence prediction. Experiments demonstrate that a well-designed compression objective can significantly enhance LLM-based text representations, outperforming models trained with token-level pretext tasks. Further improvements through contrastive learning produce a strong representation model (LLM2Comp) that outperforms contemporary LLM-based text encoders on a wide range of tasks while being more sample-efficient, requiring significantly less training data. Yeqin Zhang, Yizheng Zhao, Binxing Jiao, Daxin Jiang, Ruihang Miao, Cam-Tu Nguyen |
AAAI | 4 |
| 2024 | VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation LearningabstractAlthough speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework to integrate different modal information and leverage different resources (e.g., visual-audio pairs, audio-text pairs, unlabeled speech, and unlabeled text) to facilitate speech representation learning was not well explored. In this paper, we propose a unified cross-modal representation learning frameworkVatLM(Visual-Audio-Text Language Model). The proposedVatLMemploys a unified backbone network to model the modality-independent information and utilizes three simple modality-dependent modules to preprocess visual, speech, and text inputs. In order to integrate these three modalities into one shared semantic space,VatLMis optimized with a masked prediction task of unified tokens, given by our proposed unified tokenizer. We evaluate the pre-trainedVatLMon audio-visual related downstream tasks, including audio-visual speech recognition (AVSR), and visual speech recognition (VSR) tasks. Results show that the proposedVatLMoutperforms previous state-of-the-art models, such as the audio-visual pre-trained AV-HuBERT model, and analysis also demonstrates thatVatLMis capable of aligning different modalities into the same space. To facilitate future research, we release the code and pre-trained models athttps://aka.ms/vatlm. Qiushi Zhu, Shujie Liu 0001, Binxing Jiao, Jie Zhang 0042, Li-Rong Dai 0001, Daxin Jiang, Jinyu Li 0001, Furu Wei |
IEEE Trans. Multim. | 5 |
| 2023 | Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor WatermarkabstractWenjun Peng, Jingwei Yi, Fangzhao Wu, Shangxi Wu, Bin Bin Zhu, Lingjuan Lyu, Binxing Jiao, Tong Xu, Guangzhong Sun, Xing Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Wenjun Peng 0001, Jingwei Yi, Fangzhao Wu, Shangxi Wu, Bin B. Zhu, Lingjuan Lyu, Binxing Jiao, Tong Xu 0001, Guangzhong Sun, Xing Xie 0001 |
ACL (1) | 7 |
| 2023 | SimLM: Pre-training with Representation Bottleneck for Dense Passage RetrievalabstractLiang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liang Wang 0046, Nan Yang 0002, Xiaolong Huang 0002, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei |
ACL (1) | 4 |
| 2023 | LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval
Tao Shen 0001, Xiubo Geng, Chongyang Tao, Can Xu 0002, Xiaolong Huang 0002, Binxing Jiao, Linjun Yang, Daxin Jiang |
ICLR | 6 |
| 2023 | LED: Lexicon-Enlightened Dense Retriever for Large-Scale RetrievalabstractRetrieval models based on dense representations in semantic space have become an indispensable branch for first-stage retrieval. These retrievers benefit from surging advances in representation learning towards compressive global sequence-level embeddings. However, they are prone to overlook local salient phrases and entity mentions in texts, which usually play pivot roles in first-stage retrieval. To mitigate this weakness, we propose to make a dense retriever align a well-performing lexicon-aware representation model. The alignment is achieved by weakened knowledge distillations to enlighten the retriever via two aspects – 1) a lexicon-augmented contrastive objective to challenge the dense encoder and 2) a pair-wise rank-consistent regularization to make the dense model’s behavior incline to the other. We evaluate our model on three public benchmarks, which shows that with a comparable lexicon-aware retriever as the teacher, our proposed dense one can bring consistent and significant improvements, and even outdo its teacher. In addition, we show our lexicon-aware distillation strategies are compatible with the standard ranker distillation, which can further lift state-of-the-art performance.1 Kai Zhang 0033, Chongyang Tao, Tao Shen 0001, Can Xu 0002, Xiubo Geng, Binxing Jiao, Daxin Jiang |
WWW | 6 |
| 2022 | Effective and Efficient Query-aware Snippet Extraction for Web SearchabstractQuery-aware webpage snippet extraction is widely used in search engines to help users better understand the content of the returned webpages before clicking.Although important, it is very rarely studied.In this paper, we propose an effective query-aware webpage snippet extraction method named DeepQSE, aiming to select a few sentences which can best summarize the webpage content in the context of input query.DeepQSE first learns query-aware sentence representations for each sentence to capture the fine-grained relevance between query and sentence, and then learns document-aware query-sentence relevance representations for snippet extraction.Since the query and each sentence are jointly modeled in DeepQSE, its online inference may be slow.Thus, we further propose an efficient version of DeepQSE, named Efficient-DeepQSE, which can significantly improve the inference speed of Deep-QSE without affecting its performance.The core idea of Efficient-DeepQSE is to decompose the query-aware snippet extraction task into two stages, i.e., a coarse-grained candidate sentence selection stage where sentence representations can be cached, and a fine-grained relevance modeling stage.Experiments on two real-world datasets validate the effectiveness and efficiency of our methods. Jingwei Yi, Fangzhao Wu, Chuhan Wu, Binxing Jiao, Guangzhong Sun, Xing Xie 0001 |
EMNLP | 5 |
| 2021 | xMoCo: Cross Momentum Contrastive Learning for Open-Domain Question AnsweringabstractNan Yang, Furu Wei, Binxing Jiao, Daxing Jiang, Linjun Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Nan Yang 0002, Furu Wei, Binxing Jiao, Daxing Jiang, Linjun Yang |
ACL/IJCNLP (1) | 3 |
| 2012 | Visually Summarizing Web Pages Through Internal and External ImagesabstractVisually summarizing web pages is an attractive approach that provides users an effective and friendly interface to identify desired contents at a first glance for search and re-finding tasks. Using dominant images in web pages is generally reliable for this purpose. However, dominant images are often unavailable in many web pages. To solve this problem, we first propose a new approach to summarize those web pages without any dominant images by retrieving relevant external images from the Internet. However, relevant external images are sometimes unreliable. To take the advantages of these two kinds of images, we further propose a clustering based algorithm to select the best summarization among all of internal and external images. This algorithm leverages relevance and dominance of images as the prior information. Experimental results show that our approach achieves 0.098 and 0.082 NDCG1gain on a human labeled data set, compared with relevant external image and dominant image, respectively. Our user study also indicates that the images selected by our algorithm are useful as the summarization of web pages. Binxing Jiao, Linjun Yang, Jizheng Xu, Qi Tian 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2010 | Visual summarization of web pagesabstractVisual summarization is an attractive new scheme to summarize web pages, which can help achieve a more friendly user experience in search and re-finding tasks by allowing users quickly get the idea of what the web page is about and helping users recall the visited web page. In this paper, we perform a careful study on the recently proposed visual summarization approaches, including the thumbnail of the web page snapshot, the internal image in the web page which is representative of the content in the page, and the visual snippet which is a synthesized image based on the internal image, the title, and the logo found in the web page. Moreover, since the internal image based summarization approach hardly works when the representative internal images are unavailable, we propose a new strategy, which retrieves the representative image from the external to summarize the web page. The experimental results suggest that the various summarization approaches have respective advantages on different types of web pages. While internal images and thumbnails can provide a reliable summarization on web pages with dominant images and web pages with simple structure respectively, the external images are regarded as a useful information to complement the internal images and are demonstrated very useful in helping users understanding new web pages. The visual snippet performs well on the re-finding tasks since it incorporates the title and logo which are advantageous on identifying the visited web pages. Binxing Jiao, Linjun Yang, Jizheng Xu, Feng Wu 0001 |
SIGIR | 1 |
| 2008 | Natural Network Coding in Multi-Hop Wireless NetworksabstractCooperative communication has been intensively studied in recent years. In cooperative communication, cooperative users are grouped into clusters to form virtual MIMO channels. However, existing work mainly aims at achieving diversity gain, while generally overlooks the spatial multiplexing property of MIMO channels. In this paper, we address spatial multiplexing by introducing natural network coding protocol (NNC). In a user cooperative scenario, NNC takes advantage of the natural randomness of wireless channels. Simulation experiments show that NNC greatly reduces bit error probability and achieves about 70% throughput gain over time sharing transmission protocols. Binxing Jiao |
ICC | 3 |