VLDB 2026 Research / reviewers in the wild / expert
Binyang Li
dblp:32/7967
· DBLP profile ↗
27ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EventWeave: A Dynamic Framework for Capturing Core and Supporting Events in Dialogue SystemsabstractZhengyi Zhao, Shubo Zhang, Yiming Du, Bin Liang, Baojun Wang, Zhongyang Li, Binyang Li, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengyi Zhao 0001, Shubo Zhang, Yiming Du, Bin Liang 0004, Baojun Wang, Binyang Li, Kam-Fai Wong |
ACL (1) | 7 |
| 2026 | Guaranteeing Knowledge Integration with Joint Decoding for Retrieval-Augmented GenerationabstractZhengyi Zhao, Shubo Zhang, Zezhong Wang, Yuxi Zhang, Huimin Wang, Yutian Zhao, Yefeng Zheng, Binyang Li, Kam-Fai Wong, Xian Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengyi Zhao 0001, Shubo Zhang, Zezhong Wang 0004, Yutian Zhao, Yefeng Zheng 0001, Binyang Li, Kam-Fai Wong, Xian Wu 0001 |
ACL (1) | 8 |
| 2026 | MSCCL++: Rethinking GPU Communication Abstractions for AI InferenceabstractAI applications increasingly run on fast-evolving, heterogeneous hardware to maximize performance, but general-purpose libraries lag in supporting these features. Performance-minded programmers often build custom communication stacks that are fast but error-prone and non-portable. This paper introduces MSCCL++, a design methodology for developing high-performance, portable communication kernels. It provides (1) a low-level, performance-preserving primitive interface that exposes minimal hardware abstractions while hiding the complexities of synchronization and consistency, (2) a higher-level DSL for application developers to implement workload-specific communication algorithms, and (3) a library of efficient algorithms implementing the standard collective API, enabling adoption by users with minimal expertise. Compared to state-of-the-art baselines, MSCCL++ achieves geomean speedups of 1.7× (up to 5.4×) for collective communication and 1.2× (up to 1.38×) for AI inference workloads. MSCCL++ is in production of multiple AI services provided by Microsoft Azure, and has also been adopted by RCCL, the GPU collective communication library maintained by AMD. MSCCL++ is open source and available at https://github.com/microsoft/mscclpp. Our two years of experience with MSCCL++ suggests that its abstractions are robust, enabling support for new hardware features, such as multimem, within weeks of development. Changho Hwang, Peng Cheng 0005, Roshan Dathathri, Abhinav Jangda, Saeed Maleki, Madan Musuvathi, Olli Saarikivi, Aashaka Shah, Ziyue Yang 0002, Binyang Li, Caio Rocha, Mahdieh Ghazimirsaeed, Sreevatsa Anantharamu |
ASPLOS (2) | 10 |
| 2026 | Bridging Auxiliary Constraints to Resolve Instruction Following in Large Reasoning ModelsabstractAbstract Large Reasoning Models (LRMs) have demonstrated impressive capabilities in many tasks, yet they struggle with reliably following multiple instructions, either by failing to satisfy individual constraints or by struggling to balance competing constraints simultaneously. We formalize this challenge as the Constraint Adherence Problem (CAP). This paper introduces a novel framework that addresses CAP by representing instructions as a structured knowledge graph of constraints. Our approach, Constraint Relationship Graph Completion (CRGC), explicitly models relationships between constraints, identifies adherence challenges, and discovers “bridge constraints” that help the model better focus on and reconcile requirements. Bridge constraints act as auxiliary instructions that make primary constraints more salient and compatible. Unlike existing approaches that enhance instruction following through general training methods, CRGC specifically improves constraint satisfaction by leveraging the model’s own knowledge to create better pathways for generation. Experiments across three popular instruction following datasets demonstrate that our approach reduces constraint violations by 39% compared to standard prompting while maintaining reasoning abilities of large reasoning models. Zhengyi Zhao 0001, Shubo Zhang, Zezhong Wang 0004, Yutian Zhao, Yefeng Zheng 0001, Binyang Li, Yulan He 0001, Kam-Fai Wong, Xian Wu 0001 |
Trans. Assoc. Comput. Linguistics | 7 |
| 2025 | P2S-XR: Predictive and Scalable Scheduling for Concurrent Multi-User XR Services
Yaru Zhao 0001, Shou-lu Hou, Binyang Li, Yakun Huang |
IEEE Big Data | 3 |
| 2025 | T2: An Adaptive Test-Time Scaling Strategy for Contextual Question AnsweringabstractZhengyi Zhao, Shubo Zhang, Zezhong Wang, Huimin Wang, Yutian Zhao, Bin Liang, Yefeng Zheng, Binyang Li, Kam-Fai Wong, Xian Wu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhengyi Zhao 0001, Shubo Zhang, Zezhong Wang 0004, Yutian Zhao, Bin Liang 0004, Yefeng Zheng 0001, Binyang Li, Kam-Fai Wong, Xian Wu 0001 |
EMNLP | 8 |
| 2025 | MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language ModelsabstractZhengyi Zhao, Shubo Zhang, Yuxi Zhang, Yanxi Zhao, Yifan Zhang, Zezhong Wang, Huimin Wang, Yutian Zhao, Bin Liang, Yefeng Zheng, Binyang Li, Kam-Fai Wong, Xian Wu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhengyi Zhao 0001, Shubo Zhang, Yanxi Zhao, Yifan Zhang 0004, Zezhong Wang 0004, Yutian Zhao, Bin Liang 0004, Yefeng Zheng 0001, Binyang Li, Kam-Fai Wong, Xian Wu 0001 |
EMNLP | 11 |
| 2024 | MUCH: A Multimodal Corpus Construction for Conversational Humor Recognition Based on Chinese SitcomabstractConversational humor is the key to capturing dialogue semantics and dialogue comprehension, which is usually generated in multiple modalities, such as linguistic rhetoric (textual modality), exaggerated facial expressions or movements (visual modality), and quirky intonation (acoustic modality). However, existing multimodal corpora for conversation humor are coarse-grained, and the modality is insufficient to support the conversational humor recognition task. This paper designed an annotation scheme for multimodal humor datasets, and constructed a corpus based on a Chinese sitcom for conversational humor recognition, named MUCH. The MUCH corpus consists of 34,804 utterances in total, and 7,079 of them are humorous. We employed both unimodal and multimodal methods to test our MUCH corpus. Experimental results showed that the multimodal approach could achieve 75.94% in terms of F1-score and surpassed the performance of most unimodal methods, which demonstrated that the MUCH corpus was effective for multimodal humor recognition tasks. Wenbo Shang 0001, Xueyao Zhang, Binyang Li |
LREC/COLING | 4 |
| 2024 | UNeXt++: A Serial-Parallel Hybrid UNeXt for Rapid Medical Image Segmentation
Juelin Wang, Yunteng Deng, Binyang Li, Junlin Hu 0001 |
ICPR (28) | 4 |
| 2024 | TGAN: Temporal-Aware Graph Attention Network for Early Rumor Detection in Social Media
Shubo Zhang, Zhengyi Zhao 0001, Binyang Li, Kam-Fai Wong |
NLPCC (4) | 4 |
| 2023 | SiloD: A Co-design of Caching and Scheduling for Deep Learning ClustersabstractDeep learning training on cloud platforms usually follows the tradition of the separation of storage and computing. The training executes on a compute cluster equipped with GPUs/TPUs while reading data from a separate cluster hosting the storage service. To alleviate the potential bottleneck, a training cluster usually leverages its local storage as a cache to reduce the remote IO from the storage cluster. However, existing deep learning schedulers do not manage storage resources thus fail to consider the diverse caching effects across different training jobs. This could degrade scheduling quality significantly. Zhenhua Han, Zhi Yang 0001, Quanlu Zhang, Mingxia Li, Fan Yang 0024, Qianxi Zhang, Binyang Li, Yuqing Yang 0001, Lili Qiu, Lidong Zhou |
EuroSys | 8 |
| 2022 | Social Bot-Aware Graph Neural Network for Early Rumor DetectionabstractEarly rumor detection is a key challenging task to prevent rumors from spreading widely. Sociological research shows that social bots’ behavior in the early stage has become the main reason for rumors’ wide spread. However, current models do not explicitly distinguish genuine users from social bots, and their failure in identifying rumors timely. Therefore, this paper aims at early rumor detection by accounting for social bots’ behavior, and presents a Social Bot-Aware Graph Neural Network, named SBAG. SBAG firstly pre-trains a multi-layer perception network to capture social bot features, and then constructs multiple graph neural networks by embedding the features to model the early propagation of posts, which is further used to detect rumors. Extensive experiments on three benchmark datasets show that SBAG achieves significant improvements against the baselines and also identifies rumors within 3 hours while maintaining more than 90% accuracy. Zhen Huang 0006, Zhilong Lv, Xiaoyun Han, Binyang Li, Menglong Lu, Dongsheng Li 0001 |
COLING | 4 |
| 2022 | SIFTER: A Framework for Robust Rumor DetectionabstractWith the development of online social media, the fabrication and dissemination of rumors are much easier than before. As a result, automatic rumor detection becomes more urgent for internet governors, and has received great interest from the AI community. Even though many initial successes have been achieved in terms of detection accuracy and timeliness, existing rumor detection methods are still not robust due to thefrequent domain shiftproblem and theinconsistent propagationproblem.Frequent domain shiftrefers that rumors’ topic changes frequently on online social media, so the trained model may expire quickly in real applications.Inconsistent propagationis another practical challenge, referring to the fact that the spread of rumors is often accompanied by misleading comments. Therefore, a vulnerable model may output inconsistent predictions for the same instance. In this paper, we propose a new framework to improve the robustness of existing rumor detection methods, named SIFTER (SubjectiveInFormaTionEnhancedReinforcement learning). To address thefrequent domain shift, SIFTER employs multi-task learning to introduce external knowledge that can explicitly describe rumors, and thus makes models more adaptive across domains. In particular, SIFTER involves a multi-task learning module to capture subjective information, which is found to be discriminative in rumors’ propagations. To address theinconsistent propagation, SIFTER employs reinforcement learning to realize a new training schema tailed for rumor detection, namely sequential training, which can reduce the attendance of noisy comments in the feature extraction process. Experimental results on two benchmark datasets, Twitter and Weibo, validate the effectiveness of SIFTER. After migrating existing methods to SIFTER, their accuracy score is improved up to 9.4% in cross-domain inference, and their reverse ratio decreases by nearly 4% in continuous prediction. Menglong Lu, Zhen Huang 0006, Binyang Li, Zheng Qin 0002, Dongsheng Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | CHIME: Cross-passage Hierarchical Memory Network for Generative Review Question AnsweringabstractWe introduce CHIME, a cross-passage hierarchical memory network for question answering (QA) via text generation.It extends XLNet (Yang et al., 2019) introducing an auxiliary memory module consisting of two components: the context memory collecting cross-passage evidence, and the answer memory working as a buffer continually refining the generated answers.Empirically, we show the efficacy of the proposed architecture in the multi-passage generative QA, outperforming the state-of-the-art baselines with better syntactically well-formed answers and increased precision in addressing the questions of the AmazonQA review dataset.An additional qualitative analysis revealed the interpretability introduced by the memory module. Junru Lu, Gabriele Pergola, Lin Gui 0003, Binyang Li, Yulan He 0001 |
COLING | 4 |
| 2019 | Context-aware Embedding for Targeted Aspect-based Sentiment AnalysisabstractAttention-based neural models were employed to detect the different aspects and sentiment polarities of the same target in targeted aspectbased sentiment analysis (TABSA).However, existing methods do not specifically pre-train reasonable embeddings for targets and aspects in TABSA.This may result in targets or aspects having the same vector representations in different contexts and losing the contextdependent information.To address this problem, we propose a novel method to refine the embeddings of targets and aspects.Such pivotal embedding refinement utilizes a sparse coefficient vector to adjust the embeddings of target and aspect from the context.Hence the embeddings of targets and aspects can be refined from the highly correlative words instead of using context-independent or randomly initialized vectors.Experiment results on two benchmark datasets show that our approach yields the state-of-the-art performance in TABSA task. Bin Liang 0004, Jiachen Du, Ruifeng Xu 0001, Binyang Li, Hejiao Huang |
ACL (1) | 4 |
| 2019 | Adversarial Training Based Cross-Lingual Emotion Cause Extraction
Hongyu Yan, Qinghong Gao, Jiachen Du, Binyang Li, Ruifeng Xu 0001 |
CICLing (2) | 4 |
| 2019 | An Environment-Aware Market Strategy for Data Allocation and Dynamic Migration in Cloud DatabaseabstractCurrently, a cloud database is employed to serve on-line query-intensive applications. It inevitably happens that some cloud data nodes storing hot records are facing high frequent query requests while others are rarely visited or even idle. Therefore, how data are dynamically allocated and migrated at runtime has significant impact on query load distribution and system performance. Existing system adopt centralized approaches, and they face two main challenges: (1) Query load on individual node cannot be always balancing even if the data are fairly distributed; (2) For each node, the dynamic changes of configuration resources cannot be captured during the runtime. To this end, this paper presents an environment-aware market strategy based system, named e-MARS, for reasonable data migration to achieve query load balance in cloud database. In e-MARS, cloud database is modeled as a cloudDB market, while data nodes are regarded as intelligent traders and the query load as commodity. Each trader is aware of its local environmental re-sources, such as computing capacity, disk volume, based on which the trader itself decides how to trade the query load and migrates the corresponding data. In this way the cloudDB market will achieve equilibrium. Experiments are conducted on the real communication data, and e-MARS significantly enhances the efficiency. Compared with HBase Balancer, more than 65% improvement is achieved in terms of query response time. Tengjiao Wang 0003, Binyang Li, Wei Chen 0021, Jinzhong Niu, Kam-Fai Wong |
ICDE | 2 |
| 2018 | The UIR Uncertainty Corpus for Chinese: Annotating Chinese Microblog Corpus for Uncertainty Identification from Social Media
Binyang Li, Ruifeng Xu 0001, Tengjiao Wang 0003, Kam-Fai Wong |
LREC | 1 |
| 2017 | PFSI.sw: A programming framework for sea ice model algorithms based on Sunway many-core processorabstractSea ice model is a typical high performance computing problem. CPU and GPU based parallel method has been proposed to accelerate the simulation process, but it is still hard to meet the large-scale calculation demand due to the compute-intensive nature of the model. Sunway TaihuLight supercomputer use the SW26010 processor as its computing unit and achieves high performance for large-scale scientific computing. In this paper we present a programming framework (PFSI.sw) for sea ice model algorithms based on Sunway many-core processor. Based on this framework, programmer can exploit the parallelism of existing sea ice model algorithms and achieve good performance. Several strategies are introduced to this framework, data dividing, data transfer as well as the load balance are the main aspects we currently concerned. This framework has been implemented and tested with two sea ice model algorithms by using real world dataset on Sunway many-core processors. The experiment demonstrates comparable performance to the traditional parallel implementation on Sunway many-core processor and our framework improves the performance up to 40%. Binyang Li, Bo Li 0098, Depei Qian 0001 |
ASAP | 1 |
| 2015 | The Role of Physical Location in Our Online Social Networks
Jia Zhu 0003, Gabriel Pui Cheong Fung, Kam-Fai Wong, Binyang Li, Zhixu Li, Haoye Dong |
WAIM | 4 |
| 2014 | CLUSM: An Unsupervised Model for Microblog Sentiment Analysis Incorporating Link Information
Gaoyan Ou, Wei Chen 0021, Binyang Li, Tengjiao Wang 0003, Dongqing Yang, Kam-Fai Wong |
DASFAA (1) | 3 |
| 2014 | Exploiting Community Emotion for Microblog Event DetectionabstractMicroblog has become a major platform for information about real-world events.Automatically discovering realworld events from microblog has attracted the attention of many researchers.However, most of existing work ignore the importance of emotion information for event detection.We argue that people's emotional reactions immediately reflect the occurring of real-world events and should be important for event detection.In this study, we focus on the problem of communityrelated event detection by community emotions.To address the problem, we propose a novel framework which include the following three key components: microblog emotion classification, community emotion aggregation and community emotion burst detection.We evaluate our approach on real microblog data sets.Experimental results demonstrate the effectiveness of the proposed framework. Gaoyan Ou, Wei Chen 0021, Tengjiao Wang 0003, Zhongyu Wei, Binyang Li, Dongqing Yang, Kam-Fai Wong |
EMNLP | 5 |
| 2014 | The CUHK Discourse TreeBank for Chinese: Annotating Explicit Discourse Connectives for the Chinese TreeBank
Lanjun Zhou, Binyang Li, Zhongyu Wei, Kam-Fai Wong |
LREC | 2 |
| 2013 | Is Twitter A Better Corpus for Measuring Sentiment Similarity?abstractExtensive experiments have validated the effectiveness of the corpus-based method for classifying the word's sentiment polarity.However, no work is done for comparing different corpora in the polarity classification task.Nowadays, Twitter has aggregated huge amount of data that are full of people's sentiments.In this paper, we empirically evaluate the performance of different corpora in sentiment similarity measurement, which is the fundamental task for word polarity classification.Experiment results show that the Twitter data can achieve a much better performance than the Google, Web1T and Wikipedia based methods. Shi Feng 0001, Binyang Li, Daling Wang, Ge Yu 0001, Kam-Fai Wong |
EMNLP | 3 |
| 2011 | Unsupervised Discovery of Discourse Relations for Eliminating Intra-sentence Polarity Ambiguities
Lanjun Zhou, Binyang Li, Wei Gao 0001, Zhongyu Wei, Kam-Fai Wong |
EMNLP | 2 |
| 2010 | A Unified Graph Model for Sentence-Based Opinion Retrieval
Binyang Li, Lanjun Zhou, Shi Feng 0001, Kam-Fai Wong |
ACL | 1 |
| 2010 | Summarizing and Extracting Online Public Opinion from Blog Search Results
Shi Feng 0001, Daling Wang, Ge Yu 0001, Binyang Li, Kam-Fai Wong |
DASFAA (1) | 4 |