Binxing Jiao

dblp:78/2418 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-4710-0095ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Information retrieval · 100%
Artificial intelligence
7 papers
Representation and self-supervised learning · 71% Language models and text generation · 11% Speech recognition and synthesis · 11%
Network and information security
1 paper
Security and privacy of machine learning · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval models
2.442023
LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval · WWW 2023
LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval · ICLR 2023
Effective and Efficient Query-aware Snippet Extraction for Web Search · EMNLP 2022
Information retrieval › retrieval models › neural retrieval
dense retrieval
1.832023
LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval · WWW 2023
LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval · ICLR 2023
xMoCo: Cross Momentum Contrastive Learning for Open-Domain Question Answering · ACL/IJCNLP (1) 2021
Machine learning › Representation and self-supervised learning
contrastive learning
1.222026
Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026
LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval · ICLR 2023
Machine learning › Representation and self-supervised learning › text embedding
text representation learning
1.012026
Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026
Machine learning › Representation and self-supervised learning › text embedding › sentence embedding
unsupervised sentence embeddings
1.012026
Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026
Program synthesis and code generation
code agent
1.012026
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging · AAAI 2026
Machine learning › Representation and self-supervised learning
pre-training
0.922023
LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval · ICLR 2023
SimLM: Pre-training with Representation Bottleneck for Dense Passage Retrieval · ACL (1) 2023
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining
0.812024
VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation Learning · IEEE Trans. Multim. 2024
Natural language and speech › Speech recognition and synthesis
speech representation learning
0.812024
VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation Learning · IEEE Trans. Multim. 2024
Information retrieval › document retrieval › passage retrieval
dense passage retrieval
0.712023
SimLM: Pre-training with Representation Bottleneck for Dense Passage Retrieval · ACL (1) 2023
Information retrieval › retrieval models
knowledge distillation for retrieval
0.712023
LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval · WWW 2023
Security and privacy of machine learning
model intellectual property protection
0.712023
Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark · ACL (1) 2023
Security and privacy of machine learning › model intellectual property protection
model watermarking
0.712023
Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark · ACL (1) 2023
Information retrieval
search engines
0.622022
Effective and Efficient Query-aware Snippet Extraction for Web Search · EMNLP 2022
Visually Summarizing Web Pages Through Internal and External Images · IEEE Trans. Multim. 2012
Information retrieval › text summarization
query-biased snippet generation
0.612022
Effective and Efficient Query-aware Snippet Extraction for Web Search · EMNLP 2022
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.512021
xMoCo: Cross Momentum Contrastive Learning for Open-Domain Question Answering · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation › large language model
large language model adaptation
0.312026
Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation · AAAI 2026
Natural language and speech › Language models and text generation
LLM agents
0.312026
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging · AAAI 2026
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked prediction
0.212024
VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation Learning · IEEE Trans. Multim. 2024
Natural language and speech › Language models and text generation
large language model
0.212023
Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark · ACL (1) 2023
Information retrieval
image retrieval
0.112012
Visually Summarizing Web Pages Through Internal and External Images · IEEE Trans. Multim. 2012
Information retrieval › search interfaces
search result presentation
0.112012
Visually Summarizing Web Pages Through Internal and External Images · IEEE Trans. Multim. 2012
Information retrieval › interactive information retrieval › search tasks
re-finding
0.112010
Visual summarization of web pages · SIGIR 2010
Multimedia analysis and retrieval
image retrieval
0.012010
Visual summarization of web pages · SIGIR 2010

Methods — techniques the papers use, named apart from their topics

contrastive learning · 3.0pre-training · 2.6benchmark evaluation · 2.0embedding-based watermark · 1.3backdoor watermarking · 1.3masked next-token prediction · 1.0context compression · 1.0unified tokenizer · 0.8masked prediction · 0.8cross-modal representation learning · 0.8representation bottleneck · 0.7rank-consistent regularization · 0.7knowledge distillation · 0.7two-stage coarse-to-fine inference · 0.6neural sentence representation · 0.6momentum contrastive learning · 0.5cross momentum · 0.5user study · 0.4
YearPublicationVenuePosition
2026 GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
abstract
Beyond scratch coding, exploiting large-scale code repositories (e.g., GitHub) for practical tasks is vital in real-world software development, yet current benchmarks rarely evaluate code agents in such authentic, workflow-driven scenarios. To bridge this gap, we introduce GitTaskBench, a benchmark designed to systematically assess this capability via 54 realistic tasks across 7 modalities and 7 domains. Each task pairs a relevant repository with an automated, human-curated evaluation harness specifying practical success criteria. Beyond measuring execution and task success, we also propose the alpha-value metric to quantify the economic benefit of agent performance, which integrates task success rates, token cost, and average developer salaries. Experiments across three state-of-the-art agent frameworks with multiple advanced LLMs show that leveraging code repositories for complex task solving remains challenging: even the best-performing system, OpenHands+Claude 3.7, solves only 48.15% of tasks. Error analysis attributes over half of failures to seemingly mundane yet critical steps like environment setup and dependency resolution, highlighting the need for more robust workflow management and increased timeout preparedness. By releasing GitTaskBench, we aim to drive progress and attention toward repository-aware code reasoning, execution, and deployment---moving agents closer to solving complex, end-to-end real-world tasks.
Ziyi Ni, Huacan Wang, Shuo Lu, Wang You, Zhenheng Tang, Sen Hu 0005, Bo Li 0117, Binxing Jiao, Daxin Jiang, Yuntao Du 0001
AAAI11
2026 Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation
abstract
Text representation plays a critical role in tasks like clustering, retrieval, and other downstream applications. With the emergence of large language models (LLMs), there is increasing interest in harnessing their capabilities for this purpose. However, most of the LLMs are inherently causal and optimized for next-token prediction, making them suboptimal for producing holistic representations. To address this, recent studies introduced pretext tasks to adapt LLMs for text representation. Most of these tasks, however, rely on token-level prediction objectives, such as the masked next-token prediction (MNTP) used in LLM2Vec. In this work, we explore the untapped potential of context compression as a pretext task for unsupervised adaptation of LLMs. During compression pre-training, the model learns to generate compact memory tokens, which substitute the whole context for downstream sequence prediction. Experiments demonstrate that a well-designed compression objective can significantly enhance LLM-based text representations, outperforming models trained with token-level pretext tasks. Further improvements through contrastive learning produce a strong representation model (LLM2Comp) that outperforms contemporary LLM-based text encoders on a wide range of tasks while being more sample-efficient, requiring significantly less training data.
Yeqin Zhang, Yizheng Zhao, Binxing Jiao, Daxin Jiang, Ruihang Miao, Cam-Tu Nguyen
AAAI4
2024 VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation Learning
abstract
Although speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework to integrate different modal information and leverage different resources (e.g., visual-audio pairs, audio-text pairs, unlabeled speech, and unlabeled text) to facilitate speech representation learning was not well explored. In this paper, we propose a unified cross-modal representation learning frameworkVatLM(Visual-Audio-Text Language Model). The proposedVatLMemploys a unified backbone network to model the modality-independent information and utilizes three simple modality-dependent modules to preprocess visual, speech, and text inputs. In order to integrate these three modalities into one shared semantic space,VatLMis optimized with a masked prediction task of unified tokens, given by our proposed unified tokenizer. We evaluate the pre-trainedVatLMon audio-visual related downstream tasks, including audio-visual speech recognition (AVSR), and visual speech recognition (VSR) tasks. Results show that the proposedVatLMoutperforms previous state-of-the-art models, such as the audio-visual pre-trained AV-HuBERT model, and analysis also demonstrates thatVatLMis capable of aligning different modalities into the same space. To facilitate future research, we release the code and pre-trained models athttps://aka.ms/vatlm.
Qiushi Zhu, Shujie Liu 0001, Binxing Jiao, Jie Zhang 0042, Li-Rong Dai 0001, Daxin Jiang, Jinyu Li 0001, Furu Wei
IEEE Trans. Multim.5
2023 Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark
abstract
Wenjun Peng, Jingwei Yi, Fangzhao Wu, Shangxi Wu, Bin Bin Zhu, Lingjuan Lyu, Binxing Jiao, Tong Xu, Guangzhong Sun, Xing Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Wenjun Peng 0001, Jingwei Yi, Fangzhao Wu, Shangxi Wu, Bin B. Zhu, Lingjuan Lyu, Binxing Jiao, Tong Xu 0001, Guangzhong Sun, Xing Xie 0001
ACL (1)7
2023 SimLM: Pre-training with Representation Bottleneck for Dense Passage Retrieval
abstract
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liang Wang 0046, Nan Yang 0002, Xiaolong Huang 0002, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei
ACL (1)4
2023 LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval
Tao Shen 0001, Xiubo Geng, Chongyang Tao, Can Xu 0002, Xiaolong Huang 0002, Binxing Jiao, Linjun Yang, Daxin Jiang
ICLR6
2023 LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval
abstract
Retrieval models based on dense representations in semantic space have become an indispensable branch for first-stage retrieval. These retrievers benefit from surging advances in representation learning towards compressive global sequence-level embeddings. However, they are prone to overlook local salient phrases and entity mentions in texts, which usually play pivot roles in first-stage retrieval. To mitigate this weakness, we propose to make a dense retriever align a well-performing lexicon-aware representation model. The alignment is achieved by weakened knowledge distillations to enlighten the retriever via two aspects – 1) a lexicon-augmented contrastive objective to challenge the dense encoder and 2) a pair-wise rank-consistent regularization to make the dense model’s behavior incline to the other. We evaluate our model on three public benchmarks, which shows that with a comparable lexicon-aware retriever as the teacher, our proposed dense one can bring consistent and significant improvements, and even outdo its teacher. In addition, we show our lexicon-aware distillation strategies are compatible with the standard ranker distillation, which can further lift state-of-the-art performance.1
Kai Zhang 0033, Chongyang Tao, Tao Shen 0001, Can Xu 0002, Xiubo Geng, Binxing Jiao, Daxin Jiang
WWW6
2022 Effective and Efficient Query-aware Snippet Extraction for Web Search
abstract
Query-aware webpage snippet extraction is widely used in search engines to help users better understand the content of the returned webpages before clicking.Although important, it is very rarely studied.In this paper, we propose an effective query-aware webpage snippet extraction method named DeepQSE, aiming to select a few sentences which can best summarize the webpage content in the context of input query.DeepQSE first learns query-aware sentence representations for each sentence to capture the fine-grained relevance between query and sentence, and then learns document-aware query-sentence relevance representations for snippet extraction.Since the query and each sentence are jointly modeled in DeepQSE, its online inference may be slow.Thus, we further propose an efficient version of DeepQSE, named Efficient-DeepQSE, which can significantly improve the inference speed of Deep-QSE without affecting its performance.The core idea of Efficient-DeepQSE is to decompose the query-aware snippet extraction task into two stages, i.e., a coarse-grained candidate sentence selection stage where sentence representations can be cached, and a fine-grained relevance modeling stage.Experiments on two real-world datasets validate the effectiveness and efficiency of our methods.
Jingwei Yi, Fangzhao Wu, Chuhan Wu, Binxing Jiao, Guangzhong Sun, Xing Xie 0001
EMNLP5
2021 xMoCo: Cross Momentum Contrastive Learning for Open-Domain Question Answering
abstract
Nan Yang, Furu Wei, Binxing Jiao, Daxing Jiang, Linjun Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Nan Yang 0002, Furu Wei, Binxing Jiao, Daxing Jiang, Linjun Yang
ACL/IJCNLP (1)3
2012 Visually Summarizing Web Pages Through Internal and External Images
abstract
Visually summarizing web pages is an attractive approach that provides users an effective and friendly interface to identify desired contents at a first glance for search and re-finding tasks. Using dominant images in web pages is generally reliable for this purpose. However, dominant images are often unavailable in many web pages. To solve this problem, we first propose a new approach to summarize those web pages without any dominant images by retrieving relevant external images from the Internet. However, relevant external images are sometimes unreliable. To take the advantages of these two kinds of images, we further propose a clustering based algorithm to select the best summarization among all of internal and external images. This algorithm leverages relevance and dominance of images as the prior information. Experimental results show that our approach achieves 0.098 and 0.082 NDCG1gain on a human labeled data set, compared with relevant external image and dominant image, respectively. Our user study also indicates that the images selected by our algorithm are useful as the summarization of web pages.
Binxing Jiao, Linjun Yang, Jizheng Xu, Qi Tian 0001, Feng Wu 0001
IEEE Trans. Multim.1
2010 Visual summarization of web pages
abstract
Visual summarization is an attractive new scheme to summarize web pages, which can help achieve a more friendly user experience in search and re-finding tasks by allowing users quickly get the idea of what the web page is about and helping users recall the visited web page. In this paper, we perform a careful study on the recently proposed visual summarization approaches, including the thumbnail of the web page snapshot, the internal image in the web page which is representative of the content in the page, and the visual snippet which is a synthesized image based on the internal image, the title, and the logo found in the web page. Moreover, since the internal image based summarization approach hardly works when the representative internal images are unavailable, we propose a new strategy, which retrieves the representative image from the external to summarize the web page. The experimental results suggest that the various summarization approaches have respective advantages on different types of web pages. While internal images and thumbnails can provide a reliable summarization on web pages with dominant images and web pages with simple structure respectively, the external images are regarded as a useful information to complement the internal images and are demonstrated very useful in helping users understanding new web pages. The visual snippet performs well on the re-finding tasks since it incorporates the title and logo which are advantageous on identifying the visited web pages.
Binxing Jiao, Linjun Yang, Jizheng Xu, Feng Wu 0001
SIGIR1
2008 Natural Network Coding in Multi-Hop Wireless Networks
abstract
Cooperative communication has been intensively studied in recent years. In cooperative communication, cooperative users are grouped into clusters to form virtual MIMO channels. However, existing work mainly aims at achieving diversity gain, while generally overlooks the spatial multiplexing property of MIMO channels. In this paper, we address spatial multiplexing by introducing natural network coding protocol (NNC). In a user cooperative scenario, NNC takes advantage of the natural randomness of wireless channels. Simulation experiments show that NNC greatly reduces bit error probability and achieves about 70% throughput gain over time sharing transmission protocols.
Binxing Jiao
ICC3