EDBT 2026 Demo / reviewers in the wild / expert
Ming Yan 0008
dblp:51/5332-8
· DBLP profile ↗
88ranked-venue papers
8as first author
73since 2021 · last 2026
0000-0003-4959-8878ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 60 · 2 first-author · 52 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 7 first-author · 35 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProFuser: Progressive Fusion of Large Language ModelsabstractWhile fusing the capacities and advantages of various large language models offers a pathway to construct more powerful and versatile models, a fundamental challenge is to properly select advantageous model during training. Existing fusion methods primarily focus on the training mode that uses cross entropy on ground truth in a teacher-forcing setup to measure a model's advantage, which may provide limited insight towards model advantage. In this paper, we introduce a novel approach that enhances the fusion process by incorporating both the training and inference modes. Our method evaluates model advantage not only through cross entropy during training but also by considering inference outputs, providing a more comprehensive assessment. To combine the two modes effectively, we introduce ProFuser to progressively transition from inference mode to training mode. To validate ProFuser's effectiveness, we fused three models, including Vicuna-7B-v1.5, Llama-2-7B-Chat, and MPT-7B-8K-Chat, and demonstrated the improved performance in knowledge, reasoning, and safety compared to baseline methods. Tianyuan Shi, Fanqi Wan, Canbin Huang, Xiaojun Quan, Chenliang Li 0003, Ming Yan 0008, Ji Zhang 0011, Minhua Huang 0002, Wu Kai |
AAAI | 6 |
| 2026 | Efficient and Effective In-context Demonstration Selection with CoresetabstractIn-context learning (ICL) has emerged as a powerful paradigm for Large Visual Language Models (LVLMs), enabling them to leverage a few examples directly from input contexts. However, the effectiveness of this approach is heavily reliant on the selection of demonstrations, a process that is NP-hard. Traditional strategies, including random, similarity-based sampling and infoscore-based sampling, often lead to inefficiencies or suboptimal performance, struggling to balance both efficiency and effectiveness in demonstration selection. In this paper, we propose a novel demonstration selection framework named Coreset-based Dual Retrieval (CoDR). We show that samples within a diverse subset achieve a higher expected mutual information. To implement this, we introduce a cluster-pruning method to construct a diverse coreset that aligns more effectively with the query while maintaining diversity. Additionally, we develop a dual retrieval mechanism that enhances the selection process by achieving global demonstration selection while preserving efficiency. Experimental results demonstrate that our method significantly improves the ICL performance compared to the existing strategies, providing a robust solution for effective and efficient demonstration selection. Zihua Wang, Jiarui Wang 0002, Haiyang Xu 0001, Ming Yan 0008, Fei Huang 0002, Xu Yang 0021, Xiu-Shen Wei, Siya Mi, Yu Zhang 0004 |
AAAI | 4 |
| 2026 | AgentOCR: Reimagining Agent History via Optical Self-CompressionabstractLang Feng, Fuchao Yang, Feng Chen, Xin Cheng, Haiyang Xu, Zhenglin Wan, Ming Yan, Bo An. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lang Feng 0002, Fuchao Yang, Xin Cheng 0007, Haiyang Xu 0001, Zhenglin Wan, Ming Yan 0008, Bo An 0001 |
ACL (1) | 7 |
| 2026 | Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement LearningabstractXuanyu Lei, Chenliang Li, Yuning Wu, Kaiming Liu, Weizhou Shen, Peng Li, Ming Yan, Fei Huang, Ya-Qin Zhang, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xuanyu Lei, Chenliang Li 0003, Yuning Wu 0001, Kaiming Liu, Weizhou Shen, Peng Li 0030, Ming Yan 0008, Fei Huang 0002, Ya-Qin Zhang, Yang Liu 0005 |
ACL (1) | 7 |
| 2026 | Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent CollaborationabstractWith the rapid advancement of post-training techniques for reasoning and information seeking, large language models (LLMs) can incorporate a large quantity of retrieved knowledge to solve complex tasks. However, the limited context window of LLMs obstructs scaling the amount of external knowledge input, prohibiting further improvement. Existing context window extension methods inevitably cause information loss. LLM-based multi-agent methods emerge as a new paradigm to handle massive input in a distributional manner, where we identify two core bottlenecks in existing agent orchestration designs. In this work, we develop a multi-agent framework, \textbf{\ExtAgents}, to overcome the bottlenecks and enable better scalability in inference-time knowledge integration without longer-context training. Benchmarked with our enhanced multi-hop question answering test, \textbf{$\boldsymbol{\infty}$Bench+}, and other public test sets including long survey generation, \ExtAgents significantly enhances the performance over existing non-training methods with the same amount of external knowledge input, regardless of whether it falls \emph{within or exceeds the context window}. Moreover, the method maintains efficiency due to high parallelism. We believe further study in the coordination of LLM agents on increasing external knowledge input could benefit real-world applications. Zhennan Wan, Peng Li 0030, Ming Yan 0008, Fei Huang 0002, Yang Liu 0005 |
ACL (1) | 4 |
| 2026 | Experience-driven Multi-turn Reinforcement Learning for GUI AgentsabstractZhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen, Haiyang Xu, Ziwei Zheng, Weiming Lu, Ming Yan, Fei Huang, Jun Xiao, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengxi Lu, Jiabo Ye, Fei Tang 0005, Yongliang Shen 0001, Haiyang Xu 0001, Ziwei Zheng, Weiming Lu 0001, Ming Yan 0008, Fei Huang 0002, Jun Xiao 0001, Yueting Zhuang |
ACL (1) | 8 |
| 2026 | MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment GroundingabstractFuwen Luo, Shengfeng Lou, Chi Chen, Ziyue Wang, Chenliang Li, Weizhou Shen, Jiyue Guo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Fuwen Luo, Shengfeng Lou, Chi Chen 0005, Ziyue Wang 0002, Chenliang Li 0003, Weizhou Shen, Jiyue Guo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Yang Liu 0005 |
ACL (1) | 9 |
| 2026 | L-CLIPScore: A Lightweight Embedding-Based Captioning Metric for Evaluating and TrainingabstractWe propose a novel embedding-based captioning metric termed asL-CLIPScorethat can be used for efficiently evaluating caption quality and training captioning model. L-CLIPScore is calculated from a lightweight CLIP (L-CLIP), which is a dual-encoder architecture compressed and distilled from CLIP. To compress, we apply two powerful techniques which are weight multiplexing and matrix decomposition for reducing the parameters of encoders and word embedding matrix, respectively. To distill, we design a novel multi-modal Similarity Regulator (SR) loss to transfer more vision-language alignment knowledge. Specifically, SR loss amplifies the multi-modal embedding similarity if the given image-text pair is matched and diminishes the similarity if the pair is non-matched. By compressing and distilling by this novel SR loss, our L-CLIP achieves comparable multi-modal alignment ability to the original CLIP while it requires fewer computation resources and running time. We carry out exhaustive experiments to validate the efficiency and effectiveness of L-CLIPScore when using it as the judge to evaluate caption quality. We also discover that when using L-CLIPScore as the supervisor to train the captioning model, it should be mixed up by an n-gram-based metric and meanwhile analyze why using L-CLIPScore only will cause fail training Yingzhe Peng, Xu Yang 0021, Ruoxi Cheng, Haiyang Xu 0001, Ming Yan 0008, Fei Huang 0002 |
IEEE Trans. Multim. | 6 |
| 2026 | Adaptively Clustering Neighbor Elements for Image-Text GenerationabstractWe propose a novel Transformer-based image-to-text generation model termed asACFthat adaptively clusters vision patches into object regions and language words into phrases to implicitly learn object-phrase alignments for better visual-text coherence. To achieve this, we design a novel self-attention layer that applies self-attention over the elements in a local cluster window instead of the whole sequence. The window size is softly decided by a clustering matrix that is calculated by the current input data and thus this process is adaptive. By stacking these revised self-attention layers to construct ACF, the small clusters in the lower layers can be grouped into a bigger cluster, e.g., vision/language. ACF clusters small objects/phrases into bigger ones. In this gradual clustering process, a parsing tree is generated which embeds the hierarchical knowledge of the input sequence. As a result, by using ACF to build the vision encoder and language decoder, the hierarchical object-phrase alignments are embedded and then transferred from vision to language domains in two popular image-to-text tasks: Image captioning and Visual Question Answering. The experiment results demonstrate the effectiveness of ACF, which outperforms most SOTA captioning and VQA models and achieves comparable scores compared with some large-scale pre-trained models. Zihua Wang, Xu Yang 0021, Haiyang Xu 0001, Hanwang Zhang, Ming Yan 0008, Fei Huang 0002, Yu Zhang 0004 |
IEEE Trans. Multim. | 5 |
| 2025 | A Training-free LLM-based Approach to General Chinese Character Error CorrectionabstractChinese spelling correction (CSC) is a crucial task that aims to correct character errors in Chinese text.While conventional CSC focuses on character substitution errors caused by mistyping, two other common types of character errors, missing and redundant characters, have received less attention.These errors are often excluded from CSC datasets during the annotation process or ignored during evaluation, even when they have been annotated.This issue limits the practicality of the CSC task.To address this issue, we introduce the task of General Chinese Character Error Correction (C2EC), which focuses on all three types of character errors.We construct a high-quality C2EC benchmark by combining and manually verifying data from CCTC and Lemon datasets.We extend the training-free prompt-free CSC method to C2EC by using Levenshtein distance for handling length changes and leveraging an additional prompt-based large language model (LLM) to improve performance.Experiments show that our method enables a 14B-parameter LLM to be on par with models nearly 50 times larger on both conventional CSC and C2EC tasks, without any fine-tuning. Houquan Zhou 0001, Bo Zhang 0071, Zhenghua Li, Ming Yan 0008, Min Zhang 0005 |
ACL (1) | 4 |
| 2025 | mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document UnderstandingabstractAnwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye, Ming Yan, Ji Zhang, Qin Jin, Fei Huang, Jingren Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Anwen Hu, Haiyang Xu 0001, Jiabo Ye, Ming Yan 0008, Ji Zhang 0011, Qin Jin, Fei Huang 0002, Jingren Zhou 0001 |
ACL (1) | 5 |
| 2025 | Mutual-Taught for Co-adapting Policy and Reward ModelsabstractTianyuan Shi, Canbin Huang, Fanqi Wan, Longguang Zhong, Ziyi Yang, Weizhou Shen, Xiaojun Quan, Ming Yan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tianyuan Shi, Canbin Huang, Fanqi Wan, Longguang Zhong, Weizhou Shen, Xiaojun Quan, Ming Yan 0008 |
ACL (1) | 8 |
| 2025 | AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient OptimizationabstractRecently, model merging methods have demonstrated powerful strengths in combining abilities on various tasks from multiple Large Language Models (LLMs). While previous model merging methods mainly focus on merging homogeneous models with identical architecture, they meet challenges when dealing with Multimodal Large Language Models (MLLMs) with inherent heterogeneous property, including differences in model architecture and the asymmetry in the parameter space. In this work, we propose AdaMMS1, a novel model merging method tailored for heterogeneous MLLMs. Our method tackles the challenges in three steps: mapping, merging and searching. Specifically, we first design mapping function between models to apply model merging on MLLMs with different architecture. Then we apply linear interpolation on model weights to actively adapt the asymmetry in the heterogeneous MLLMs. Finally in the hyper-parameter searching step, we propose an unsupervised hyper-parameter selection method for model merging. As the first model merging method capable of merging heterogeneous MLLMs without labeled data, extensive experiments on various model combinations demonstrated that AdaMMS outperforms previous model merging methods on various vision-language benchmarks.2 Yiyang Du, Xiaochen Wang 0002, Chi Chen 0005, Jiabo Ye, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Zhifang Sui, Maosong Sun 0001, Yang Liu 0005 |
CVPR | 7 |
| 2025 | SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference OptimizationabstractAs language models continue to scale, Large Language Models (LLMs) have exhibited emerging capabilities in In-Context Learning (ICL), enabling them to solve language tasks by prefixing a few in-context demonstrations (ICDs) as context. Inspired by these advancements, researchers have extended these techniques to develop Large Multimodal Models (LMMs) with ICL capabilities. However, existing LMMs face a critical issue: they often fail to effectively leverage the visual context in multimodal demonstrations and instead simply follow textual patterns. This indicates that LMMs do not achieve effective alignment between multimodal demonstrations and model outputs. To address this problem, we propose Symbol Demonstration Direct Preference Optimization (SymDPO). Specifically, SymDPO aims to break the traditional paradigm of constructing multimodal demonstrations by using random symbols to replace text answers within instances. This forces the model to carefully understand the demonstration images and establish a relationship between the images and the symbols to answer questions correctly. We validate the effectiveness of this method on multiple benchmarks, demonstrating that with SymDPO, LMMs can more effectively understand the multimodal context within examples and utilize this knowledge to answer questions better. Code is available at https://github.com/APiaoG/SymDPO. Hongrui Jia, Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Mengfan Dong, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
CVPR | 6 |
| 2025 | Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document UnderstandingabstractMultimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand.As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities.However, due to different types of annotation noise in training, current MLLMs often face conflicts between perception and cognition.Taking a document VQA task (cognition) as an example, an MLLM might generate answers that do not match the corresponding visual content identified by its OCR (perception).This conflict suggests that the MLLM might struggle to establish an intrinsic connection between the information it "sees" and what it "understands".Such conflicts challenge the intuitive notion that cognition is consistent with perception, hindering the performance and explainability of MLLMs.In this paper, we define the conflicts between cognition and perception as Cognition and Perception (C&P) knowledge conflicts, a form of multimodal knowledge conflicts, and systematically assess them with a focus on document understanding.Our analysis reveals that even GPT-4o, a leading MLLM, achieves only 75.26% C&P consistency.To mitigate the C&P knowledge conflicts, we propose a novel method called Multimodal Knowledge Consistency Fine-tuning.Our method reduces C&P knowledge conflicts across all tested MLLMs and enhances their performance in both cognitive and perceptual tasks. Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Qi Zheng 0002, Ming Yan 0008, Jiajun Bu |
EMNLP | 8 |
| 2025 | mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language ModelsabstractMulti-modal Large Language Models have demonstrated remarkable capabilities in executing instructions for a variety of single-image tasks. Despite this progress, significant challenges remain in modeling long image sequences. In this work, we introduce the versatile multi-modal large language model, mPLUG-Owl3, which enhances the capability for long image-sequence understanding in scenarios that incorporate retrieved image-text knowledge, multimodal in-context examples, and lengthy videos. Specifically, we propose novel hyper attention blocks to efficiently integrate vision and language into a common language-guided semantic space, thereby facilitating the processing of extended multi-image scenarios. We conduct evaluations on 21 benchmarks that cover single/multi-image, and short/long video understanding. mPLUG-Owl3 achieves competitive performance with the state-of-the-art methods while reducing inference time and memory usage by 87.8\% and 48.5\% in average. Moreover, we propose a Distractor Resistance evaluation to assess the ability of models to maintain focus amidst distractions. mPLUG-Owl3 also demonstrates outstanding performance in distractor resistance on ultra-long visual sequence inputs. We hope that mPLUG-Owl3 can contribute to the development of more efficient and powerful multimodal large language models. Jiabo Ye, Haiyang Xu 0001, Anwen Hu, Ming Yan 0008, Qi Qian 0001, Ji Zhang 0011, Fei Huang 0002, Jingren Zhou 0001 |
ICLR | 5 |
| 2025 | Endowing Visual Reprogramming with Adversarial RobustnessabstractVisual reprogramming (VR) leverages well-developed pre-trained models (e.g., a pre-trained classifier on ImageNet) to tackle target tasks (e.g., a traffic sign recognition task), without the need for training from scratch. Despite the effectiveness of previous VR methods, all of them did not consider the adversarial robustness of reprogrammed models against adversarial attacks, which could lead to unpredictable problems in safety-crucial target tasks. In this paper, we empirically find that reprogramming pre-trained models with adversarial robustness and incorporating adversarial samples from the target task during reprogramming can both improve the adversarial robustness of reprogrammed models. Furthermore, we propose a theoretically guaranteed adversarial robustness risk upper bound for VR, which validates our empirical findings and could provide a theoretical foundation for future research. Extensive experiments demonstrate that by adopting the strategies revealed in our empirical findings, the adversarial robustness of reprogrammed models can be enhanced. Xin Cheng 0007, Haiyang Xu 0001, Ming Yan 0008, Tao Xiang 0001, Feng Liu 0003, Lei Feng 0006 |
ICLR | 4 |
| 2025 | Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement LearningabstractOnline fine-tuning vision-language model (VLM) agents with reinforcement learning (RL) has shown promise for equipping agents with multi-step, goal-oriented capabilities in dynamic environments. However, their open-ended textual action space and non-end-to-end nature of action generation present significant challenges to effective online exploration in RL, e.g., explosion of the exploration space. We propose a novel online fine-tuning method, Counterfactual Soft Reinforcement Learning (CoSo), better suited to the textual output space of VLM agents. Compared to prior methods that assign uniform uncertainty to all tokens, CoSo leverages counterfactual reasoning to dynamically assess the causal influence of individual tokens on post-processed actions. By prioritizing the exploration of action-critical tokens while reducing the impact of semantically redundant or low-impact tokens, CoSo enables a more targeted and efficient online rollout process. We provide theoretical analysis proving CoSo's convergence and policy improvement guarantees, and extensive empirical evaluations supporting CoSo's effectiveness. Our results across a diverse set of agent tasks, including Android device control, card gaming, and embodied AI, highlight its remarkable ability to enhance exploration efficiency and deliver consistent performance gains. The code is available at https://github.com/langfengQ/CoSo. Lang Feng 0002, Weihao Tan, Zhiyi Lyu, Longtao Zheng, Haiyang Xu 0001, Ming Yan 0008, Fei Huang 0002, Bo An 0001 |
ICML | 6 |
| 2025 | Exploiting Presentative Feature Distributions for Parameter-Efficient Continual Learning of Large Language ModelsabstractEndowing large language models (LLMs) with continual learning (CL) capacities is practically important, which enables them to dynamically acquire new knowledge over time. Although many effective methods have been proposed for CL of LLMs, they did not consider online scenarios, thereby sharing a common problem: information leakage (IL), where the task-related information of learned tasks is accessed or reused again. IL not only imposes potential risks on data privacy protection but also significantly hinders the deployment of LLMs in real-world scenarios. To avoid IL while maintaining outstanding CL performance, we propose a novel CL method for LLMs, which first characterizes a parameter-efficient fine-tuning (PEFT) block by a presentative feature distribution, and then dynamically selects the appropriate PEFT blocks for each instance based on its similarity with the presentative feature distributions. Extensive experiments validate the effectiveness of our method on the CL of LLM, showcasing its potential to enhance both privacy and adaptability in practical applications. Xin Cheng 0007, Jiabo Ye, Haiyang Xu 0001, Ming Yan 0008, Ji Zhang 0011, Feng Liu 0003, Fei Huang 0002, Lei Feng 0006 |
ICML | 4 |
| 2025 | VLM-R³: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-ThoughtabstractRecently, reasoning-based MLLMs have achieved a degree of success in generating long-form textual reasoning chains. However, they still struggle with complex tasks that necessitate dynamic and iterative focusing on and revisiting of visual regions to achieve precise grounding of textual reasoning in visual evidence. We introduce VLM-R³ (Visual Language Model with Region Recognition, Reasoning, and Refinement ), a framework that equips an MLLM with the ability to (i) decide when additional visual evidence is needed, (ii) determine where to ground within the image, and (iii) seamlessly weave the relevant sub-image content back into an interleaved chain-of-thought. The core of our method is \textbf{Region-Conditioned Reinforcement Policy Optimization (R-GRPO)}, a training paradigm that rewards the model for selecting informative regions, formulating appropriate transformations (e.g. crop, zoom), and integrating the resulting visual context into subsequent reasoning steps. To bootstrap this policy, we compile a modest but carefully curated Visuo-Lingual Interleaved Rationale (VLIR) corpus that provides step-level supervision on region selection and textual justification. Extensive experiments on MathVista, ScienceQA, and other benchmarks show that VLM-R$^3$ sets a new state of the art in zero-shot and few-shot settings, with the largest gains appearing on questions demanding subtle spatial reasoning or fine-grained visual cue extraction. Chaoya Jiang, Yongrui Heng, Wei Ye 0004, Haiyang Xu 0001, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
NeurIPS | 5 |
| 2025 | Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI AutomationabstractIn recent years, Multimodal Large Language Models (MLLMs) have been extensively utilized for multimodal reasoning tasks, including Graphical User Interface (GUI) automation. Unlike general offline multimodal tasks, GUI automation is executed in online interactive environments, necessitating step-by-step decision-making based on the real-time status of the environment. This task has a lower tolerance for decision-making errors at each step, as any mistakes may cumulatively disrupt the process and potentially lead to irreversible outcomes like deletions or payments. To address these issues, we introduce a pre-operative critic mechanism that provides effective feedback prior to the actual execution, by reasoning about the potential outcome and correctness of actions. Specifically, we propose a Suggestion-aware Group Relative Policy Optimization (S-GRPO) strategy to construct our pre-operative critic model GUI-Critic-R1, incorporating a novel suggestion reward to enhance the reliability of the model's feedback. Furthermore, we develop a reasoning-bootstrapping based data collection pipeline to create a GUI-Critic-Train and a GUI-Critic-Test, filling existing gaps in GUI critic data. Static experiments on the GUI-Critic-Test across both mobile and web domains reveal that our GUI-Critic-R1 offers significant advantages in critic accuracy compared to current MLLMs. Dynamic evaluation on GUI automation benchmark further highlights the effectiveness and superiority of our model, as evidenced by improved success rates and operational efficiency. The code is available at https://github.com/X-PLUG/MobileAgent/tree/main/GUI-Critic-R1. Yuyang Wanyan, Haiyang Xu 0001, Junyang Wang 0001, Jiabo Ye, Yutong Kou, Ming Yan 0008, Fei Huang 0002, Xiaoshan Yang, Weiming Dong, Changsheng Xu |
NeurIPS | 8 |
| 2025 | WritingBench: A Comprehensive Benchmark for Generative WritingabstractRecent advancements in large language models (LLMs) have significantly enhanced text generation capabilities, yet evaluating their performance in generative writing remains a challenge. Existing benchmarks primarily focus on generic text generation or limited in writing tasks, failing to capture the diverse requirements of high-quality written contents across various domains. To bridge this gap, we present WritingBench, a comprehensive benchmark designed to evaluate LLMs across 6 core writing domains and 100 subdomains. We further propose a query-dependent evaluation framework that empowers LLMs to dynamically generate instance-specific assessment criteria. This framework is complemented by a fine-tuned critic model for criteria-aware scoring, enabling evaluations in style, format and length. The framework's validity is further demonstrated by its data curation capability, which enables a 7B-parameter model to outperform the performance of GPT-4o in writing. We open-source the benchmark, along with evaluation tools and modular framework components, to advance the development of LLMs in writing. Yuning Wu 0001, Jiahao Mei, Ming Yan 0008, Chenliang Li 0003, Shaopeng Lai, Yuran Ren, Ji Zhang 0011, Mengyue Wu, Qin Jin, Fei Huang 0002 |
NeurIPS | 3 |
| 2024 | TiMix: Text-Aware Image Mixing for Effective Vision-Language Pre-trainingabstractSelf-supervised Multi-modal Contrastive Learning (SMCL) remarkably advances modern Vision-Language Pre-training (VLP) models by aligning visual and linguistic modalities. Due to noises in web-harvested text-image pairs, however, scaling up training data volume in SMCL presents considerable obstacles in terms of computational cost and data inefficiency. To improve data efficiency in VLP, we propose Text-aware Image Mixing (TiMix), which integrates mix-based data augmentation techniques into SMCL, yielding significant performance improvements without significantly increasing computational overhead. We provide a theoretical analysis of TiMix from a mutual information (MI) perspective, showing that mixed data samples for cross-modal contrastive learning implicitly serve as a regularizer for the contrastive loss. The experimental results demonstrate that TiMix exhibits a comparable performance on downstream tasks, even with a reduced amount of training data and shorter training time, when benchmarked against existing methods. This work empirically and theoretically demonstrates the potential of data mixing for data-efficient and computationally viable VLP, benefiting broader VLP model adoption in practical scenarios. Our code is available on https://github.com/chaoyajiang/TiMiX/tree/main. Chaoya Jiang, Wei Ye 0004, Haiyang Xu 0001, Qinghao Ye, Ming Yan 0008, Ji Zhang 0011, Shikun Zhang |
AAAI | 5 |
| 2024 | Model Composition for Multimodal Large Language ModelsabstractChi Chen, Yiyang Du, Zheng Fang, Ziyue Wang, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chi Chen 0005, Yiyang Du, Ziyue Wang 0002, Fuwen Luo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 7 |
| 2024 | Browse and Concentrate: Comprehending Multimodal Content via Prior-LLM Context FusionabstractZiyue Wang, Chi Chen, Yiqi Zhu, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ziyue Wang 0002, Chi Chen 0005, Yiqi Zhu, Fuwen Luo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 6 |
| 2024 | SiTunes: A Situational Music Recommendation Dataset with Physiological and Psychological SignalsabstractWith an increasing number of music tracks available online, music recommender systems have become popular and ubiquitous. Previous research indicates that people’s preferences, especially in music, dynamically change with various factors, such as surrounding situations and emotional status. However, few existing public recommendation datasets contain such situation or emotion information. Therefore, we constructed SiTunes, a situational music recommendation dataset with rich physiological and psychological signals. We collected the data through a three-stage user study, including: (1) recorded users’ inherent music preference in a lab setting (Stage 1), (2) recorded physiological and environmental situations by smart wristband devices in users’ daily life, and provided psychological and rating feedback for music recommended by traditional recommenders (Stage 2) and (3) by situation-aware recommenders (Stage 3). The experiments were conducted with strict privacy concerns and ethical approval. The dataset contains over 2000 listening logs from 30 users on over 300 music tracks. SiTunes serves as a valuable resource for future studies on situational recommenders and user understanding in recommendation. The dataset is available at https://github.com/JiayuLi-997/SiTunes_dataset/. Vadim Grigorev, Jiayu Li 0001, Weizhi Ma, Zhiyu He 0001, Min Zhang 0006, Yiqun Liu 0001, Ming Yan 0008, Ji Zhang 0011 |
CHIIR | 7 |
| 2024 | Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-trainingabstractIn vision-language pre-training (VLP), masked image modeling (MIM) has recently been introduced for fine-grained cross-modal alignment. However, in most existing methods, the reconstruction targets for MIM lack high-level semantics, and text is not sufficiently involved in masked modeling. These two drawbacks limit the effect of MIM in facilitating cross-modal semantic alignment. In this work, we propose a semantics-enhanced cross-modal MIM framework (SemMIM) for vision-language representation learning. Specifically, to provide more semantically meaningful supervision for MIM, we propose a local semantics enhancing approach, which harvest high-level semantics from global image features via self-supervised agreement learning and transfer them to local patch encodings by sharing the encoding space. Moreover, to achieve deep involvement of text during the entire MIM process, we propose a text-guided masking strategy and devise an efficient way of injecting textual information in both masked modeling and reconstruction target acquisition. Experimental results validate that our method improves the effectiveness of the MIM task in facilitating cross-modal semantic alignment. Compared to previous VLP models with similar model size and data scale, our SemMIM model achieves state-of-the-art or competitive performance on multiple downstream vision-language tasks. Yaya Shi, Haiyang Xu 0001, Chunfeng Yuan, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Bing Li 0001, Weiming Hu 0004 |
LREC/COLING | 7 |
| 2024 | Unifying Latent and Lexicon Representations for Effective Video-Text RetrievalabstractIn video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representations for videos and texts. However, they face challenges in capturing fine-grained semantic concepts. In this work, we propose the UNIFY framework, which learns lexicon representations to capture fine-grained semantics and combines the strengths of latent and lexicon representations for video-text retrieval. Specifically, we map videos and texts into a pre-defined lexicon space, where each dimension corresponds to a semantic concept. A two-stage semantics grounding approach is proposed to activate semantically relevant dimensions and suppress irrelevant dimensions. The learned lexicon representations can thus reflect fine-grained semantics of videos and texts. Furthermore, to leverage the complementarity between latent and lexicon representations, we propose a unified learning scheme to facilitate mutual learning via structure sharing and self-distillation. Experimental results show our UNIFY framework largely outperforms previous video-text retrieval methods, with 4.8% and 8.2% Recall@1 improvement on MSR-VTT and DiDeMo respectively. Yaya Shi, Haiyang Xu 0001, Chunfeng Yuan, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Bing Li 0001, Weiming Hu 0004 |
LREC/COLING | 7 |
| 2024 | Hallucination Augmented Contrastive Learning for Multimodal Large Language ModelabstractMulti-modal large language models (MLLMs) have been shown to efficiently integrate natural language with visual information to handle multi-modal tasks. However, MLLMs still face a fundamental limitation of hallucinations, where they tend to generate erroneous or fabricated information. In this paper, we address hallucinations in MLLMs from a novel perspective of representation learning. We first analyzed the representation distribution of textual and visual tokens in MLLM, revealing two important findings: 1) there is a significant gap between textual and visual representations, indicating unsatisfactory cross-modal representation alignment; 2) representations of texts that contain and do not contain hallucinations are entangled, making it challenging to distinguish them. These two observations inspire us with a simple yet effective method to mitigate hallucinations. Specifically, we introduce contrastive learning into MLLMs and use text with hallucination as hard negative examples, naturally bringing representations of non-hallucinative text and visual samples closer while pushing way representations of non-hallucinating and hallucinative text. We evaluate our method quantitatively and qualitatively, showing its effectiveness in reducing hallucination occurrences and improving performance across multiple benchmarks. On the MMhal-Bench benchmark, our method obtains a 34.66% /29.5% improvement over the baseline MiniGPT-4/LLaVA. Our code is available on https://github.com/X-PLUG/mPLUG-HalOwl/tree/main/hacl. Chaoya Jiang, Haiyang Xu 0001, Mengfan Dong, Wei Ye 0004, Ming Yan 0008, Qinghao Ye, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
CVPR | 6 |
| 2024 | mPLUG-OwI2: Revolutionizing Multi-modal Large Language Model with Modality CollaborationabstractMulti-modal Large Language Models (MLLMs) have demonstrated impressive instruction abilities across various open-ended tasks. However, previous methods primarily fo-cus on enhancing multi-modal capabilities. In this work, we introduce a versatile multi-modal large language model, mPLUG-Owl2, which effectively leverages modality collab-oration to improve performance in both text and multi-modal tasks. mPLUG-Owl2 utilizes a modularized network design, with the language decoder acting as a universal interface for managing different modalities. Specifically, mPLUG-Owl2 incorporates shared functional modules to facilitate modal-ity collaboration and introduces a modality-adaptive module that preserves modality-specific features. Extensive experi-ments reveal that mPLUG-Owl2 is capable of generalizing both text tasks and multi-modal tasks and achieving state-of-the-art performances with a single generic model. Notably, mPLUG-Owl2 is the first MLLM model that demonstrates the modality collaboration phenomenon in both pure-text and multi-modal scenarios, setting a pioneering path in the development of future multi-modal foundation models. Qinghao Ye, Haiyang Xu 0001, Jiabo Ye, Ming Yan 0008, Anwen Hu, Qi Qian 0001, Ji Zhang 0011, Fei Huang 0002 |
CVPR | 4 |
| 2024 | MIBench: Evaluating Multimodal Large Language Models over Multiple ImagesabstractHaowei Liu, Xi Zhang, Haiyang Xu, Yaya Shi, Chaoya Jiang, Ming Yan, Ji Zhang, Fei Huang, Chunfeng Yuan, Bing Li, Weiming Hu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Haiyang Xu 0001, Yaya Shi, Chaoya Jiang, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004 |
EMNLP | 6 |
| 2024 | Small LLMs Are Weak Tool Learners: A Multi-LLM AgentabstractLarge Language Model (LLM) agents significantly extend the capabilities of standalone LLMs, empowering them to interact with external tools (e.g., APIs, functions) and complete various tasks in a self-directed fashion.The challenge of tool use demands that LLMs not only understand user queries and generate answers accurately but also excel in task planning, tool invocation, and result summarization.While traditional works focus on training a single LLM with all these capabilities, performance limitations become apparent, particularly with smaller models.To overcome these challenges, we propose a novel approach that decomposes the aforementioned capabilities into a planner, caller, and summarizer.Each component is implemented by a single LLM that focuses on a specific capability and collaborates with others to accomplish the task.This modular framework facilitates individual updates and the potential use of smaller LLMs for building each capability.To effectively train this framework, we introduce a two-stage training paradigm.First, we fine-tune a backbone LLM on the entire dataset without discriminating sub-tasks, providing the model with a comprehensive understanding of the task.Second, the fine-tuned LLM is used to instantiate the planner, caller, and summarizer respectively, which are continually fine-tuned on respective sub-tasks.Evaluation across various tool-use benchmarks illustrates that our proposed multi-LLM framework surpasses the traditional single-LLM approach, highlighting its efficacy and advantages in tool learning. Weizhou Shen, Chenliang Li 0003, Hongzhan Chen, Ming Yan 0008, Xiaojun Quan, Hehong Chen, Ji Zhang 0011, Fei Huang 0002 |
EMNLP | 4 |
| 2024 | TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token MergingabstractCharts are important for presenting and explaining complex data relationships.Recently, multimodal large language models (MLLMs) have shown remarkable capabilities in chart understanding.However, the sheer size of these models limits their use in resource-constrained environments.In this paper, we present Tiny-Chart, an efficient MLLM for chart understanding with only 3B parameters.TinyChart overcomes two key challenges in efficient chart understanding: (1) reduce the burden of learning numerical computations through Programof-Thoughts (PoT) learning, which trains the model to generate Python programs for numerical calculations, and (2) reduce lengthy vision feature sequences through Vision Token Merging, which gradually merges most similar vision tokens.Extensive experiments demonstrate that our 3B TinyChart achieves SOTA performance on various chart understanding benchmarks including ChartQA, Chart-to-Text, Chart-to-Table, OpenCQA, and ChartX.It outperforms several chart-understanding MLLMs with up to 13B parameters, and close-sourced MLLM GPT-4V on ChartQA, with higher throughput during inference due to a smaller model scale and more efficient vision encoding. Anwen Hu, Haiyang Xu 0001, Ming Yan 0008, Yichen Xu 0003, Qin Jin, Ji Zhang 0011, Fei Huang 0002 |
EMNLP | 4 |
| 2024 | Two-Stage Information Bottleneck For Temporal Language GroundingabstractExisting cross-modal fusion methods for temporal language grounding suffer from issues where the generated cross-modal embeddings are affected by noise in their unimodal representations, thus hindering the expression of interactions between language query and the target video segment. Furthermore, the cross-modal representations contain many irrelevant redundancies, which compromises the quality of cross-modal features and thus interferes with accurate moment localization. To address these drawbacks, we propose a novel CrOss-modaL information-constraineD (COLD) model for temporal language grounding, aiming at learning a robust cross-modal embedding representation devoid of irrelevant redundancies and maximizing the interaction between language query and target video moment. Specifically, our model is built upon the principles of the information bottleneck and features two information-constrained modules from different perspectives: 1) the Cross-modal Highlight Information Bottleneck module is designed to maximize the mutual information between language query and target video moment; 2) the Fusion Information Bottleneck module is introduced to constrain the correlations between the cross-modal representations and the localization labels. Comprehensive experimental results on two public benchmark datasets demonstrate the superiority of the proposed model. Haoyu Tang 0002, Shuaike Zhang, Ming Yan 0008, Ji Zhang 0011, Yupeng Hu 0003, Liqiang Nie |
ICME | 3 |
| 2024 | VG-Annotator: Vision-Language Models as Query Annotators for Unsupervised Visual GroundingabstractVisual grounding focuses on localizing objects referred to by natural language queries. Existing fully and weakly supervised methods rely on a mass of language queries for training. However, collecting natural language queries corresponding to specific objects by annotators is expensive. To reduce the reliance on human-written queries, we propose a novel unsupervised visual grounding framework named VG-Annotator. Different from the existing unsupervised methods that rely on manually designed rules to link objects and language queries. The key idea of VG-Annotator lies in that vision-language pre-trained (VLP) generation models can be language query annotators. Thanks to the powerful multi-modal understanding ability implicitly learned from large-scale pre-training, we consider stimulating models to explicitly generate appropriate descriptions for specific objects in natural language. To this end, we explore a series of multi-modal instructions to indicate which object should be described. We also introduce a supervised fine-tuning process to teach the vision-language models to follow the instructions. Extensive experiments show that the proposed method obtains high-quality language queries. The visual grounding model trained with the generated queries outperforms state-of-the-art unsupervised methods on five widely used datasets. Jiabo Ye, Xiaoshan Yang, Zhenru Zhang, Anwen Hu, Ming Yan 0008, Ji Zhang 0011, Liang He 0001, Xin Lin 0001 |
ICME | 6 |
| 2024 | Breaking Barriers of System Heterogeneity: Straggler-Tolerant Multimodal Federated Learning via Knowledge Distillation
Jinqian Chen, Haoyu Tang 0002, Ming Yan 0008, Ji Zhang 0011, Yupeng Hu 0003, Liqiang Nie |
IJCAI | 4 |
| 2024 | DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
Baihan Li, Zeyu Xie, Xuenan Xu, Ming Yan 0008, Ji Zhang 0011, Kai Yu 0004, Mengyue Wu |
INTERSPEECH | 5 |
| 2024 | Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
Xuenan Xu, Pingyue Zhang, Ming Yan 0008, Ji Zhang 0011, Mengyue Wu |
INTERSPEECH | 3 |
| 2024 | mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language ModelabstractWeak diagram analysis abilities of LLMs or Multimodal LLMs greatly limit their application scenarios for scientific academic paper writing. In this work, towards a more versatile copilot for academic paper writing, we mainly focus on strengthening the multi-modal diagram analysis ability of Multimodal LLMs. By parsing Latex source files of academic papers, we carefully build a multi-modal diagram understanding dataset M-Paper. By aligning diagrams in the paper with related paragraphs, we construct professional diagram analysis samples for training and evaluation. M-Paper is the first dataset to support joint comprehension of multiple scientific diagrams, including figures and tables in the format of images or Latex codes. Besides, to better align the copilot with the user's intention, we introduce the 'outline' as the control signal, which could be directly given by the user or revised based on auto-generated ones. Comprehensive experiments with a state-of-the-art Multimodal LLM demonstrate that training on our dataset shows stronger scientific diagram understanding performance. The dataset, code, and model are publicly available at https://github.com/X-PLUG/mPLUG-DocOwl/tree/main/PaperOwl. Anwen Hu, Yaya Shi, Haiyang Xu 0001, Jiabo Ye, Qinghao Ye, Ming Yan 0008, Chenliang Li 0003, Qi Qian 0001, Ji Zhang 0011, Fei Huang 0002 |
ACM Multimedia | 6 |
| 2024 | Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language ModelsabstractLarge Vision-Language Models (LVLMs) exhibit remarkable capabilities but struggle with ''hallucinations''-inconsistencies between images and their descriptions. Previous hallucination evaluation studies on LVLMs have identified hallucinations in terms of objects, attributes, and relations but overlooked complex hallucinations that create an entire narrative around a fictional entity. In this paper, we introduce a refined taxonomy of hallucinations, featuring a new category: Event Hallucination. We then utilize advanced LLMs to generate and filter fine-grained hallucinatory data consisting of various types of hallucinations, with a particular focus on event hallucinations, laying the groundwork for integrating discriminative and generative evaluation methods within our universal evaluation framework. The proposed benchmark distinctively assesses LVLMs' ability to tackle a broad spectrum of hallucinations, making it a reliable and comprehensive tool for gauging LVLMs' efficacy in handling hallucinations. We will release our code and data. Chaoya Jiang, Hongrui Jia, Mengfan Dong, Wei Ye 0004, Haiyang Xu 0001, Ming Yan 0008, Ji Zhang 0011, Shikun Zhang |
ACM Multimedia | 6 |
| 2024 | Revisiting Unsupervised Temporal Action Localization: The Primacy of High-Quality Actionness and PseudolabelsabstractRecently, temporal action localization (TAL) methods, especially the weakly-supervised and unsupervised ones, have become a hot research topic. Existing unsupervised methods follow an iterative ''clustering and training'' strategy with diverse model designs during training stage, while they often overlook maintaining consistency between these stages, which is crucial: more accurate clustering results can reduce the noises of pseudolabels and thus enhance model training, while more robust training can in turn enrich clustering feature representation. We identify two critical challenges in unsupervised scenarios: 1. What features should the model generate for clustering? 2. Which pseudolabeled instances from clustering should be chosen for model training? After extensive explorations, we proposed a novel yet simple framework called Consistency-Oriented Progressive high actionness Learning to address these issues. For feature generation, our framework adopts a High Actionness snippet Selection (HAS) module to generate more discriminative global video features for clustering from the enhanced actionness features obtained from a designed Inner-Outer Consistency Network (IOCNet). For pseudolabel selection, we introduces a Progressive Learning With Representative Instances (PLRI) strategy to identify the most reliable and informative instances within each cluster for model training. These three modules, HAS, IOCNet, and PLRI, synergistically improve consistency in model training and clustering performance. Extensive experiments on THUMOS'14 and ActivityNet v1.2 datasets under both unsupervised and weakly-supervised settings demonstrate that our framework achieves the state-of-the-art results. Han Jiang 0012, Haoyu Tang 0002, Ming Yan 0008, Ji Zhang 0011, Yupeng Hu 0003, Jihua Zhu, Liqiang Nie |
ACM Multimedia | 3 |
| 2024 | Part-Aware Prompt Tuning for Weakly Supervised Referring Expression Grounding
Chenlin Zhao, Jiabo Ye, Yaguang Song, Ming Yan 0008, Xiaoshan Yang, Changsheng Xu |
MMM (3) | 4 |
| 2024 | Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationabstractMobile device operation tasks are increasingly becoming a popular multi-modal AI application scenario. Current Multi-modal Large Language Models (MLLMs), constrained by their training data, lack the capability to function effectively as operation assistants. Instead, MLLM-based agents, which enhance capabilities through tool invocation, are gradually being applied to this scenario. However, the two major navigation challenges in mobile device operation tasks — task progress navigation and focus content navigation — are difficult to effectively solve under the single-agent architecture of existing work. This is due to the overly long token sequences and the interleaved text-image data format, which limit performance. To address these navigation challenges effectively, we propose Mobile-Agent-v2, a multi-agent architecture for mobile device operation assistance. The architecture comprises three agents: planning agent, decision agent, and reflection agent. The planning agent condenses lengthy, interleaved image-text history operations and screens summaries into a pure-text task progress, which is then passed on to the decision agent. This reduction in context length makes it easier for decision agent to navigate the task progress. To retain focus content, we design a memory unit that updates with task progress by decision agent. Additionally, to correct erroneous operations, the reflection agent observes the outcomes of each operation and handles any mistake accordingly. Experimental results indicate that Mobile-Agent-v2 achieves over a 30% improvement in task completion compared to the single-agent architecture of Mobile-Agent. The code is open-sourced at https://github.com/X-PLUG/MobileAgent. Junyang Wang 0001, Haiyang Xu 0001, Haitao Jia, Ming Yan 0008, Weizhou Shen, Ji Zhang 0011, Fei Huang 0002, Jitao Sang 0001 |
NeurIPS | 5 |
| 2024 | MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language ModelabstractThis paper presents MaVEn, an innovative Multi-granularity Visual Encoding framework designed to enhance the capabilities of Multimodal Large Language Models (MLLMs) in multi-image reasoning. Current MLLMs primarily focus on single-image visual understanding, limiting their ability to interpret and integrate information across multiple images. MaVEn addresses this limitation by combining discrete visual symbol sequences, which abstract coarse-grained semantic concepts, with traditional continuous representation sequences that model fine-grained features. This dual approach bridges the semantic gap between visual and textual data, thereby improving the model's ability to process and interpret information from multiple images effectively. Additionally, we design a dynamic reduction mechanism by for long-sequence continuous features to enhance multi-image processing efficiency. Experimental results demonstrate that MaVEn significantly enhances MLLMs' understanding in complex multi-image scenarios, while also improving performance in single-image contexts. Chaoya Jiang, Hongrui Jia, Haiyang Xu 0001, Wei Ye 0004, Mengfan Dong, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Shikun Zhang |
NeurIPS | 6 |
| 2024 | Modeling Comparative Logical Relation with Contrastive Learning for Text Generation
Yuhao Dan, Jie Zhou 0015, Ming Yan 0008, Ji Zhang 0011, Qin Chen 0001, Liang He 0001 |
NLPCC (4) | 4 |
| 2024 | CLIP-VG: Self-Paced Curriculum Adapting of CLIP for Visual GroundingabstractVisual Grounding (VG) is a crucial topic in the field of vision and language, which involves locating a specific region described by expressions within an image. To reduce the reliance on manually labeled data, unsupervised methods have been developed to locate regions using pseudo-labels. However, the performance of existing unsupervised methods is highly dependent on the quality of pseudo-labels and these methods always encounter issues with limited diversity. In order to utilize vision and language pre-trained models to address the grounding problem, and reasonably take advantage of pseudo-labels, we propose CLIP-VG, a novel method that can conduct self-paced curriculum adapting of CLIP with pseudo-language labels. We propose a simple yet efficient end-to-end network architecture to realize the transfer of CLIP to the visual grounding. Based on the CLIP-based architecture, we further propose single-source and multi-source curriculum adapting algorithms, which can progressively find more reliable pseudo-labels to learn an optimal model, thereby achieving a balance between reliability and diversity for the pseudo-language labels. Our method outperforms the current state-of-the-art unsupervised method by a significant margin on RefCOCO/+/g datasets in both single-source and multi-source scenarios, with improvements ranging from 6.78% to 10.67% and 11.39% to 14.87%, respectively. Furthermore, our approach even outperforms existing weakly supervised methods. Linhui Xiao, Xiaoshan Yang, Ming Yan 0008, Yaowei Wang 0001, Changsheng Xu |
IEEE Trans. Multim. | 4 |
| 2024 | UniQRNet: Unifying Referring Expression Grounding and Segmentation with QRNetabstractReferring expression comprehension aims to align natural language queries with visual scenes, which requires establishing fine-grained correspondence between vision and language. This has important applications in multi-modal reasoning systems. Existing methods typically use text-agnostic visual backbones to extract features independently without considering the specific text input. However, we argue that the extracted visual features can be inconsistent with the referring expression, which hurts multi-modal understanding. To address this, we first propose Query-modulated Refinement Network (QRNet) that leverages language guidance to guide visual feature extraction. However, it only focuses on the grounding task that can only provide coarse-grained annotations in the form of bounding box coordinates. The guidance for the visual backbone is indirect, and the inconsistent issue still exists. To this end, we further propose UniQRNet, a multi-task framework over the QRNet to learn referring expression grounding and segmentation jointly. The framework introduces a multi-task head that leverages fine-grained pixel-level supervision from the segmentation task to directly guide the intermediate layers of QRNet to learn text-consistent visual features. Besides, UniQRNet also includes a loss balance strategy that allows two types of supervision signals to cooperate and optimize the model together. We conduct the most comprehensive comparison experiment covering four major datasets, ten evaluation set and three evaluation metrics used in previous work. UniQRNet outperforms previous state-of-the-art methods by a large margin on both referring comprehensive grounding (1.8%~5.09%) and segmentation tasks (0.57%~5.56%). Ablation and analysis reveal that UniQRNet can improve the consistency of visual features with text input and can bring significant performance improvement. Jiabo Ye, Ming Yan 0008, Haiyang Xu 0001, Qinghao Ye, Yaya Shi, Xiaoshan Yang, Xuwu Wang, Ji Zhang 0011, Liang He 0001, Xin Lin 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | BUS : Efficient and Effective Vision-language Pre-training with Bottom-Up Patch SummarizationabstractVision Transformer (ViT) based Vision-Language Pre-training (VLP) models have demonstrated impressive performance in various tasks. However, the lengthy visual token sequences fed into ViT can lead to training inefficiency and ineffectiveness. Existing efforts address the challenge by either bottom-level patch extraction in the ViT backbone or top-level patch abstraction outside, not balancing training efficiency and effectiveness well. Inspired by text summarization in natural language processing, we propose a Bottom-Up Patch Summarization approach named BUS, coordinating bottom-level extraction and top-level abstraction to learn a concise summary of lengthy visual token sequences efficiently. Specifically, We incorporate a Text-Semantics-Aware Patch Selector (TSPS) into the ViT backbone to perform a coarse-grained visual token extraction and then attach a flexible Transformer-based Patch Abstraction Decoder (PAD) upon the backbone for top-level visual abstraction. This bottom-up collaboration enables our BUS to yield high training efficiency while maintaining or even improving effectiveness. We evaluate our approach on various visual-language understanding and generation tasks and show competitive downstream task performance while boosting the training efficiency by 50%. Additionally, our model achieves state-of-the-art performance on many downstream tasks by increasing input image resolution without increasing computational costs over baselines. Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Bin Bi, Shikun Zhang, Fei Huang 0002, Songfang Huang |
ICCV | 6 |
| 2023 | Improved Visual Fine-tuning with Natural Language SupervisionabstractFine-tuning a visual pre-trained model can leverage the semantic information from large-scale pre-training data and mitigate the over-fitting problem on downstream vision tasks with limited training examples. While the problem of catastrophic forgetting in pre-trained backbone has been extensively studied for fine-tuning, its potential bias from the corresponding pre-training task and data, attracts less attention. In this work, we investigate this problem by demonstrating that the obtained classifier after fine-tuning will be close to that induced by the pre-trained model. To reduce the bias in the classifier effectively, we introduce a reference distribution obtained from a fixed text classifier, which can help regularize the learned vision classifier. The proposed method, Text Supervised fine-tuning (TeS), is evaluated with diverse pre-trained vision models including ResNet and ViT, and text encoders including BERT and CLIP, on 11 downstream tasks. The consistent improvement with a clear margin over distinct scenarios confirms the effectiveness of our proposal. Code is available at https://github.com/idstcv/TeS. Junyang Wang 0001, Yuanhong Xu, Juhua Hu, Ming Yan 0008, Jitao Sang 0001, Qi Qian 0001 |
ICCV | 4 |
| 2023 | Learning Trajectory-Word Alignments for Video-Language TasksabstractIn a video, an object usually appears as the trajectory, i.e., it spans over a few spatial but longer temporal patches, that contains abundant spatiotemporal contexts. However, modern Video-Language BERTs (VDL-BERTs) neglect this trajectory characteristic that they usually follow image-language BERTs (IL-BERTs) to deploy the patch-to-word (P2W) attention that may over-exploit trivial spatial contexts and neglect significant temporal contexts. To amend this, we propose a novel TW-BERT to learn Trajectory-Word alignment by a newly designed trajectory-to-word (T2W) attention for solving video-language tasks. Moreover, previous VDL-BERTs usually uniformly sample a few frames into the model while different trajectories have diverse graininess, i.e., some trajectories span longer frames and some span shorter, and using a few frames will lose certain useful temporal contexts. However, simply sampling more frames will also make pre-training infeasible due to the largely increased training burdens. To alleviate the problem, during the fine-tuning stage, we insert a novel Hierarchical Frame-Selector (HFS) module into the video encoder. HFS gradually selects the suitable frames conditioned on the text context for the later cross-modal encoder to learn better trajectory-word alignments. By the proposed T2W attention and HFS, our TW-BERT achieves SOTA performances on text-to-video retrieval tasks, and comparable performances on video question-answering tasks with some VDL-BERTs trained on much more data. The code will be available in the supplementary material. Xu Yang 0021, Zhangzikang Li, Haiyang Xu 0001, Hanwang Zhang, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Yu Zhang 0004, Fei Huang 0002, Songfang Huang |
ICCV | 7 |
| 2023 | HiTeA: Hierarchical Temporal-Aware Video-Language Pre-trainingabstractVideo-language pre-training has advanced the performance of various downstream video-language tasks. However, most previous methods directly inherit or adapt typical image-language pre-training paradigms to video-language pre-training, thus not fully exploiting the unique characteristic of video, i.e., temporal. In this paper, we propose a Hierarchical Temporal-Aware video-language pre-training framework, HiTeA, with two novel pre-training tasks for yielding temporal-aware multi-modal representation with cross-modal fine-grained temporal moment information and temporal contextual relations between video-text multi-modal pairs. First, we propose a cross-modal moment exploration task to explore moments in videos by mining the paired texts, which results in detailed video moment representation. Then, based on the learned detailed moment representations, the inherent temporal contextual relations are captured by aligning video-text pairs as a whole in different time resolutions with multi-modal temporal relation exploration task. Furthermore, we introduce the shuffling test to evaluate the temporal reliance of datasets and video-language pre-training models. We achieve state-of-the-art results on 15 well-established video-language understanding and generation tasks, especially on temporal-oriented datasets (e.g., SSv2-Template and SSv2-Label) with 8.6% and 11.1% improvement respectively. HiTeA also demonstrates strong generalization ability when directly transferred to downstream tasks in a zero-shot manner. Qinghao Ye, Guohai Xu, Ming Yan 0008, Haiyang Xu 0001, Qi Qian 0001, Ji Zhang 0011, Fei Huang 0002 |
ICCV | 3 |
| 2023 | Construction and Applications of Billion-Scale Pre-Trained Multimodal Business Knowledge GraphabstractBusiness Knowledge Graphs (KGs) are important to many enterprises today, providing factual knowledge and structured data that steer many products and make them more intelligent. Despite their promising benefits, building business KG necessitates solving prohibitive issues of deficient structure and multiple modalities. In this paper, we advance the understanding of the practical challenges related to building KG in non-trivial real-world systems. We introduce the process of building an open business knowledge graph (OpenBG) derived from a well-known enterprise, Alibaba Group. Specifically, we define a core ontology to cover various abstract products and consumption demands, with fine-grained taxonomy and multimodal facts in deployed applications. OpenBG is an open business KG of unprecedented scale: 2.6 billion triples with more than 88 million entities covering over 1 million core classes/concepts and 2,681 types of relations. We release all the open resources (OpenBG benchmarks) derived from it for the community and report experimental results of KG-centric tasks. We also run up an online competition based on OpenBG benchmarks, and has attracted thousands of teams. We further pre-train OpenBG and apply it to many KG-enhanced downstream tasks in business scenarios, demonstrating the effectiveness of billion-scale multimodal knowledge for e-commerce. All the resources with codes have been released at https://github.com/OpenBGBenchmark/OpenBG. Shumin Deng, Zhoubo Li, Ningyu Zhang 0001, Zelin Dai, Hehong Chen, Feiyu Xiong, Ming Yan 0008, Mosha Chen, Jiaoyan Chen 0001, Jeff Z. Pan, Bryan Hooi, Huajun Chen |
ICDE | 8 |
| 2023 | mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and VideoabstractRecent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for multi-modal pretraining, which can benefit from modality collaboration while addressing the problem of modality entanglement. In contrast to predominant paradigms of solely relying on sequence-to-sequence generation or encoder-based instance discrimination, mPLUG-2 introduces a multi-module composition network by sharing common universal modules for modality collaboration and disentangling different modality modules to deal with modality entanglement. It is flexible to select different modules for different understanding and generation tasks across all modalities including text, image, and video. Empirical study shows that mPLUG-2 achieves state-of-the-art or competitive results on a broad range of over 30 downstream tasks, spanning multi-modal tasks of image-text and video-text understanding and generation, and uni-modal tasks of text-only, image-only, and video-only understanding. Notably, mPLUG-2 shows new state-of-the-art results of 48.0 top-1 accuracy and 80.3 CIDEr on the challenging MSRVTT video QA and video caption tasks with a far smaller model size and data scale. It also demonstrates strong zero-shot transferability on vision-language and video-language tasks. Code and models will be released in https://github.com/X-PLUG/mPLUG-2. Haiyang Xu 0001, Qinghao Ye, Ming Yan 0008, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li 0003, Bin Bi, Qi Qian 0001, Wei Wang 0225, Guohai Xu, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Jingren Zhou 0001 |
ICML | 3 |
| 2023 | From Association to Generation: Text-only Captioning by Unsupervised Cross-modal MappingabstractWith the development of Vision-Language Pre-training Models (VLPMs) represented by CLIP and ALIGN, significant breakthroughs have been achieved for association-based visual tasks such as image classification and image-text retrieval by the zero-shot capability of CLIP without fine-tuning. However, CLIP is hard to apply to generation-based tasks. This is due to the lack of decoder architecture and pre-training tasks for generation. Although previous works have created generation capacity for CLIP through additional language models, a modality gap between the CLIP representations of different modalities and the inability of CLIP to model the offset of this gap, which results in the failure of the concept to transfer across modes. To solve the problem, we try to map images/videos to the language modality and generate captions from the language modality. In this paper, we propose the K-nearest-neighbor Cross-modality Mapping (Knight), a zero-shot method from association to generation. With vision-free unsupervised training, Knight achieves state-of-the-art performance in zero-shot methods for image captioning and video captioning. Junyang Wang 0001, Ming Yan 0008, Yi Zhang 0101, Jitao Sang 0001 |
IJCAI | 2 |
| 2023 | COPA : Efficient Vision-Language Pre-training through Collaborative Object- and Patch-Text AlignmentabstractVision-Language Pre-training (VLP) methods based on object detection enjoy the rich knowledge of fine-grained object-text alignment but at the cost of computationally expensive inference. Recent Visual-Transformer (ViT)-based approaches circumvent this issue while struggling with long visual sequences without detailed cross-modal alignment information. This paper introduces a ViT-based VLP technique that efficiently incorporates object information through a novel patch-text alignment mechanism. Specifically, we convert object-level signals into patch-level ones and devise a Patch-Text Alignment pre-training task (PTA) to learn a text-aware patch detector. By using off-the-shelf delicate object annotations in 5% training images, we jointly train PTA with other conventional VLP objectives in an end-to-end manner, bypassing the high computational cost of object detection and yielding an effective patch detector that accurately detects text-relevant patches, thus considerably reducing patch sequences and accelerating computation within the ViT backbone. Our experiments on a variety of widely-used benchmarks reveal that our method achieves a speedup of nearly 88% compared to prior VLP models while maintaining competitive or superior performance on downstream tasks with similar model size and data scale. Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Bin Bi, Shikun Zhang, Fei Huang 0002, Ji Zhang 0011 |
ACM Multimedia | 6 |
| 2023 | Learning Semantics-Grounded Vocabulary Representation for Video-Text RetrievalabstractPrevious dual-encoder pre-training methods for video-text retrieval employ contrastive learning for cross-modal alignment in a latent space. However, such learned latent spaces often result in modality gap problem [26]. In this paper, we introduce a novel SemVTR framework designed to learn semantics-grounded video-text representations in a vocabulary space, in which each dimension corresponds to a semantic concept represented by a word. The representation is obtained by grounding video and text into semantically-related dimensions with high activation values. As video-text pairs share grounded dimensions, their vocabulary representations are expected to cluster together and thus alleviate modality gap problem. So, the crux of our method lies in grounding video and text into vocabulary space. Specifically, we propose a Multi-Granularity Video Semantics Grounding approach and a Textual Semantics Preserving training strategy. The visualization illustrates that SemVTR obtains semantics-gronded vocabulary representation and also alleviates the modality gap problem. SemVTR significantly outperforms existing methods on four video-text retrieval benchmarks. Yaya Shi, Haiyang Xu 0001, Zongyang Ma, Qinghao Ye, Anwen Hu, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Zhengjun Zha |
ACM Multimedia | 7 |
| 2023 | mPLUG-Octopus: The Versatile Assistant Empowered by A Modularized End-to-End Multimodal LLMabstractInspired by the recent developments of large language models (LLMs), we propose mPLUG-Octopus, a versatile conversational assistant designed to provide users with coherent, engaging, and helpful interaction experiences in both text-only and multi-modal scenarios. Unlike traditional pipeline chatting systems, mPLUG-Octopus offers a diverse range of creative capabilities including open-domain QA, multi-turn chatting, and multi-modal creation, all built with a unified multimodal LLM without relying on any external API. With the modularized end-to-end multimodal LLM technology, mPLUG-Octopus efficiently facilitates engaging and open-domain conversation experience. It exhibits a wide range of uni/multi-modal elemental capabilities, enabling it to seamlessly communicate with users on open-domain topics and engage in multi-turn conversations. It also assists users in accomplishing various content creation and application tasks. Our conversational assistant can also be deployed on smart hardware to drive advanced AIGC applications. Qinghao Ye, Haiyang Xu 0001, Ming Yan 0008, Chenlin Zhao, Junyang Wang 0001, Xiaoshan Yang, Ji Zhang 0011, Fei Huang 0002, Jitao Sang 0001, Changsheng Xu |
ACM Multimedia | 3 |
| 2023 | Multi-modal multi-hop interaction network for dialogue response generation
Jie Zhou 0015, Rui Wang 0005, Yuanbin Wu, Ming Yan 0008, Liang He 0001, Xuanjing Huang 0001 |
Expert Syst. Appl. | 5 |
| 2023 | Attribute-Guided Collaborative Learning for Partial Person Re-IdentificationabstractPartial person re-identification (ReID) aims to solve the problem of image spatial misalignment due to occlusions or out-of-views. Despite significant progress through the introduction of additional information, such as human pose landmarks, mask maps, and spatial information, partial person ReID remains challenging due to noisy keypoints and impressionable pedestrian representations. To address these issues, we propose a unified attribute-guided collaborative learning scheme for partial person ReID. Specifically, we introduce an adaptive threshold-guided masked graph convolutional network that can dynamically remove untrustworthy edges to suppress the diffusion of noisy keypoints. Furthermore, we incorporate human attributes and devise a cyclic heterogeneous graph convolutional network to effectively fuse cross-modal pedestrian information through intra- and inter-graph interaction, resulting in robust pedestrian representations. Finally, to enhance keypoint representation learning, we design a novel part-based similarity constraint based on the axisymmetric characteristic of the human body. Extensive experiments on multiple public datasets have shown that our model achieves superior performance compared to other state-of-the-art baselines. Meng Liu 0006, Ming Yan 0008, Zan Gao 0001, Xiaojun Chang, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Achieving Human Parity on Visual Question AnsweringabstractThe Visual Question Answering (VQA) task utilizes both visual image and language analysis to answer a textual question with respect to an image. It has been a popular research topic with an increasing number of real-world applications in the last decade. This paper introduces a novel hierarchical integration of vision and language AliceMind-MMU (ALIbaba’s Collection of Encoder-decoders from Machine IntelligeNce lab of Damo academy - MultiMedia Understanding) , which leads to similar or even slightly better results than a human being does on VQA. A hierarchical framework is designed to tackle the practical problems of VQA in a cascade manner including: (1) diverse visual semantics learning for comprehensive image content understanding; (2) enhanced multi-modal pre-training with modality adaptive attention; and (3) a knowledge-guided model integration with three specialized expert modules for the complex VQA task. Treating different types of visual questions with corresponding expertise needed plays an important role in boosting the performance of our VQA architecture up to the human level. An extensive set of experiments and analysis are conducted to demonstrate the effectiveness of the new research work. Ming Yan 0008, Haiyang Xu 0001, Chenliang Li 0003, Bin Bi, Wei Wang 0225, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Luo Si, Rong Jin 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2022 | WikiDiverse: A Multimodal Entity Linking Dataset with Diversified Contextual Topics and Entity TypesabstractXuwu Wang, Junfeng Tian, Min Gui, Zhixu Li, Rui Wang, Ming Yan, Lihan Chen, Yanghua Xiao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Xuwu Wang, Min Gui, Zhixu Li, Rui Wang 0005, Ming Yan 0008, Yanghua Xiao |
ACL (1) | 6 |
| 2022 | Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual GroundingabstractVisual grounding focuses on establishing fine-grained alignment between vision and natural language, which has essential applications in multimodal reasoning systems. Existing methods use pre-trained query-agnostic visual backbones to extract visual feature maps independently without considering the query information. We argue that the visual features extracted from the visual backbones and the features really needed for multimodal reasoning are inconsistent. One reason is that there are differences between pre-training tasks and visual grounding. Moreover, since the backbones are query-agnostic, it is difficult to completely avoid the inconsistency issue by training the visual backbone end-to-end in the visual grounding framework. In this paper, we propose a Query-modulated Refinement Network (QRNet) to address the inconsistent issue by adjusting intermediate features in the visual backbone with a novel Query-aware Dynamic Attention (QD-ATT) mechanism and query-aware multiscale fusion. The QD-ATT can dynamically compute query-dependent visual attention at the spatial and channel levels of the feature maps produced by the visual backbone. We apply the QRNet to an end-to-end visual grounding framework. Extensive experiments show that the proposed method outperforms state-of-the-art methods on five widely used datasets. Our code is available at https://github.com/LukeForeverYoung/QRNet. Jiabo Ye, Ming Yan 0008, Xiaoshan Yang, Xuwu Wang, Ji Zhang 0011, Liang He 0001, Xin Lin 0001 |
CVPR | 3 |
| 2022 | PromptMNER: Prompt-Based Entity-Related Visual Clue Extraction and Integration for Multimodal Named Entity Recognition
Xuwu Wang, Min Gui, Zhixu Li, Jiabo Ye, Ming Yan 0008, Yanghua Xiao |
DASFAA (3) | 6 |
| 2022 | TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch SelectionabstractVision Transformers (ViTs) have been widely used in large-scale Vision and Language Pretraining (VLP) models.Though previous VLP works have proved the effectiveness of ViTs, they still suffer from computational efficiency brought by the long visual sequence.To tackle this problem, in this paper, we propose an efficient vision-and-language pre-training model with Text-Relevant Image Patch Selection, namely TRIPS, which reduces the visual sequence progressively with a text-guided patchselection layer in the visual backbone for efficient training and inference.The patchselection layer can dynamically compute textdependent visual attention to identify the attentive image tokens with text guidance and fuse inattentive ones in an end-to-end manner.Meanwhile, TRIPS does not introduce extra parameters to ViTs.Experimental results on a variety of popular benchmark datasets demonstrate that TRIPS gain a speedup of 40% over previous similar VLP models, yet with competitive or better downstream task performance. Chaoya Jiang, Haiyang Xu 0001, Chenliang Li 0003, Ming Yan 0008, Wei Ye 0004, Shikun Zhang, Bin Bi, Songfang Huang |
EMNLP | 4 |
| 2022 | mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connectionsabstractChenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, He Chen, Guohai Xu, Zheng Cao, Ji Zhang, Songfang Huang, Fei Huang, Jingren Zhou, Luo Si. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Chenliang Li 0003, Haiyang Xu 0001, Wei Wang 0225, Ming Yan 0008, Bin Bi, Jiabo Ye, Guohai Xu, Zheng Cao 0003, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Jingren Zhou 0001, Luo Si |
EMNLP | 5 |
| 2022 | CAT-MNER: Multimodal Named Entity Recognition with Knowledge-Refined Cross-Modal AttentionabstractMultimodal named entity recognition (MNER) aims to detect and classify named entities in multimodal scenarios. It requires bridging the gap between natural language and visual context, which presents two-fold challenges: the cross-modal alignment is diversified, and the cross-modal interaction is sometimes implicit. Existing MNER methods are vulnerable to some implicit interactions and are prone to overlook the involved significant features. To tackle this problem, we novelly propose to refine the cross-modal attention by identifying and highlighting some task-salient features. The saliency of each feature is measured according to its correlation with the expanded entity label words derived from external knowledge bases. We further propose an end-to-end Transformer-based MNER framework, which holds neater architecture yet achieves better performance than previous methods. Extensive experiments are conducted to validate the merits of our method. Moreover, our method reveals a significant advantage in data efficiency and generalization ability. Xuwu Wang, Jiabo Ye, Zhixu Li, Yong Jiang 0005, Ming Yan 0008, Ji Zhang 0011, Yanghua Xiao |
ICME | 6 |
| 2022 | DictBERT: Dictionary Description Knowledge Enhanced Language Model Pre-training via Contrastive LearningabstractAlthough pre-trained language models (PLMs) have achieved state-of-the-art performance on various natural language processing (NLP) tasks, they are shown to be lacking in knowledge when dealing with knowledge driven tasks. Despite the many efforts made for injecting knowledge into PLMs, this problem remains open. To address the challenge, we propose DictBERT, a novel approach that enhances PLMs with dictionary knowledge which is easier to acquire than knowledge graph (KG). During pre-training, we present two novel pre-training tasks to inject dictionary knowledge into PLMs via contrastive learning: dictionary entry prediction and entry description discrimination. In fine-tuning, we use the pre-trained DictBERT as a plugin knowledge base (KB) to retrieve implicit knowledge for identified entries in an input sequence, and infuse the retrieved knowledge into the input to enhance its representation via a novel extra-hop attention mechanism. We evaluate our approach on a variety of knowledge driven and language understanding tasks, including NER, relation extraction, CommonsenseQA, OpenBookQA and GLUE. Experimental results demonstrate that our model can significantly improve typical PLMs: it gains a substantial improvement of 0.5%, 2.9%, 9.0%, 7.1% and 3.3% on BERT-large respectively, and is also effective on RoBERTa-large. Qianglong Chen, Feng-Lin Li, Guohai Xu, Ming Yan 0008, Ji Zhang 0011, Yin Zhang 0006 |
IJCAI | 4 |
| 2022 | Attribute-guided Dynamic Routing Graph Network for Transductive Few-shot LearningabstractMotivated by the structured form of human cognition, attributes have been introduced in few-shot classification to learn more representative sample features. However, existing attribute-based methods usually treat the importance of different attributes as equals to conclude the sample relations, which cannot distinguish the classes with many similar attributes well. In order to address this problem, we propose an Attribute-guided Dynamic Routing Graph Network (ADRGN) to explicitly learn task-dependent attribute importance scores to help explore the sample relations in a fine-grained manner for adaptive graph-based inference. Specifically, we first leverage a CNN backbone and a transformation network to generate attribute-specific sample representations according to attribute annotations. Next, we treat the attribute-specific sample representations as visual primary capsules and employ an inter-sample routing to explore the visual diversity of each attribute in the current task. Based on the generated diversity capsules, we perform an inter-attribute routing to explore the relations between different attributes to predict the visual attribute importance scores. Meanwhile, we design an attribute semantic routing module to predict the semantic attribute importance from the semantic attribute embeddings to help the learning of the visual attribute importance prediction with a knowledge distillation strategy. Finally, we utilize the visual attribute importance scores to adaptively aggregate sample similarities computed based on the attribute-specific representations to capture the global fine-grained sample relations for message passing and graph-based inference. Experimental results on three few-shot classification benchmarks show that the proposed ADRGN obtains state-of-the-art performance. Xiaoshan Yang, Ming Yan 0008, Changsheng Xu |
ACM Multimedia | 3 |
| 2022 | X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalabstractVideo-text retrieval has been a crucial and fundamental task in multi-modal research. The development of video-text retrieval has been considerably promoted by large-scale multi-modal contrastive pre-training, which primarily focuses on coarse-grained or fine-grained contrast. However, cross-grained contrast, which is the contrast between coarse-grained representations and fine-grained representations, has rarely been explored in prior research. Compared with fine-grained or coarse-grained contrasts, cross-grained contrast calculate the correlation between coarse-grained features and each fine-grained feature, and is able to filter out the unnecessary fine-grained features guided by the coarse-grained feature during similarity calculation, thus improving the accuracy of retrieval. To this end, this paper presents a novel multi-grained contrastive model, namely X-CLIP, for video-text retrieval. However, another challenge lies in the similarity aggregation problem, which aims to aggregate fine-grained and cross-grained similarity matrices to instance-level similarity. To address this challenge, we propose the Attention Over Similarity Matrix (AOSM) module to make the model focus on the contrast between essential frames and words, thus lowering the impact of unnecessary frames and words on retrieval results. With multi-grained contrast and the proposed AOSM module, X-CLIP achieves outstanding performance on five widely-used video-text retrieval datasets, including MSR-VTT (49.3 [email protected]), MSVD (50.4 [email protected]), LSMDC (26.1 [email protected]), DiDeMo (47.8 [email protected]) and ActivityNet (46.2 [email protected]). Guohai Xu, Xiaoshuai Sun, Ming Yan 0008, Ji Zhang 0011, Rongrong Ji |
ACM Multimedia | 4 |
| 2022 | Comprehensive Relationship Reasoning for Composed Query Based Image RetrievalabstractComposed Query Based Image Retrieval (CQBIR) aims at searching images relevant to a composed query, i.e., a reference image together with a modifier text. Compared with conventional image retrieval, which takes a single image or text to retrieve desired images, CQBIR encounters more challenges as it requires not only effective semantic correspondence between the heterogeneous query and target, but also synergistic understanding of the composed query. To establish robust CQBIR model, four critical types of relational information can be included, i.e., cross-modal, intra-sample, inter-sample, and cross-sample relationships. Pioneer studies mainly exploit parts of the information, which are hard to make them enhance and complement each other. In this paper, we propose a comprehensive relationship reasoning network by fully exploring the four types of information for CQBIR, which mainly includes two key designs. First, we introduce a memory-augmented cross-modal attention module, in which the representation of the composed query is augmented by considering the cross-modal relationship between the reference image and the modification text. Second, we design a multi-scale matching strategy to optimize our network, aiming at harnessing information from the intra-sample, inter-sample, and cross-sample relationships. To the best of our knowledge, this is the first work to fully explore the four pieces of relationships in a unified deep model for CQBIR. Comprehensive experimental results on five standard benchmarks demonstrate that the proposed method performs favorably against state-of-the-art models. Feifei Zhang 0001, Ming Yan 0008, Ji Zhang 0011, Changsheng Xu |
ACM Multimedia | 2 |
| 2021 | A Unified Pretraining Framework for Passage Ranking and ExpansionabstractPretrained language models have recently advanced a wide range of natural language processing tasks. Nowadays, the application of pretrained language models to IR tasks has also achieved impressive results. Typical methods either directly apply a pretrained model to improve the re-ranking stage, or use it to conduct passage expansion and term weighting for first-stage retrieval. We observe that the passage ranking and passage expansion tasks share certain inherent relations, and can benefit from each other. Therefore, in this paper, we propose a general pretraining framework to enhance both tasks with Unified Encoder-Decoder networks (UED). The overall ranking framework consists of two parts in a cascade manner: (1) passage expansion with a pretraining-based query generation method; (2) re-ranking of passage candidates from a traditional retrieval method with a pretrained transformer encoder. Both the two parts are based on the same pretrained UED model, where we jointly train the passage ranking and query generation tasks for further improving the full ranking pipeline. An extensive set of experiments have been conducted on two large-scale passage retrieval datasets to demonstrate the state-of-the-art results of the proposed framework in both the first-stage retrieval and the final re-ranking. In addition, we successfully deploy the framework to our online production system, which can stably serve industrial applications with a request volume of up to 100 QPS in less than 300ms. Ming Yan 0008, Chenliang Li 0003, Bin Bi, Wei Wang 0225, Songfang Huang |
AAAI | 1 |
| 2021 | StructuralLM: Structural Pre-training for Form UnderstandingabstractChenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chenliang Li 0003, Bin Bi, Ming Yan 0008, Wei Wang 0225, Songfang Huang, Fei Huang 0002, Luo Si |
ACL/IJCNLP (1) | 3 |
| 2021 | E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual LearningabstractHaiyang Xu, Ming Yan, Chenliang Li, Bin Bi, Songfang Huang, Wenming Xiao, Fei Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Haiyang Xu 0001, Ming Yan 0008, Chenliang Li 0003, Bin Bi, Songfang Huang, Wenming Xiao, Fei Huang 0002 |
ACL/IJCNLP (1) | 2 |
| 2020 | Generating Well-Formed Answers by Machine Reading with Stochastic Selector Networks
Bin Bi, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Jiangnan Xia, Chenliang Li 0003 |
AAAI | 3 |
| 2020 | PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned GenerationabstractSelf-supervised pre-training, such as BERT (Devlin et al., 2018), MASS (Song et al., 2019) and BART (Lewis et al., 2019), has emerged as a powerful technique for natural language understanding and generation.Existing pre-training techniques employ autoencoding and/or autoregressive objectives to train Transformer-based models by recovering original word tokens from corrupted text with some masked tokens.The training goals of existing techniques are often inconsistent with the goals of many language generation tasks, such as generative question answering and conversational response generation, for producing new text given context.This work presents PALM with a novel scheme that jointly pre-trains an autoencoding and autoregressive language model on a large unlabeled corpus, specifically designed for generating new text conditioned on context.The new scheme alleviates the mismatch introduced by the existing denoising scheme between pre-training and fine-tuning where generation is more than reconstructing original text.An extensive set of experiments show that PALM achieves new state-of-theart results on a variety of language generation benchmarks covering generative question answering (Rank 1 on the official MARCO leaderboard), abstractive summarization on CNN/DailyMail as well as Gigaword, question generation on SQuAD, and conversational response generation on Cornell Movie Dialogues. Bin Bi, Chenliang Li 0003, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Songfang Huang, Fei Huang 0002, Luo Si |
EMNLP (1) | 4 |
| 2020 | StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
Wei Wang 0225, Bin Bi, Ming Yan 0008, Chen Wu 0006, Jiangnan Xia, Zuyi Bao, Liwei Peng, Luo Si |
ICLR | 3 |
| 2019 | A Deep Cascade Model for Multi-Document Reading ComprehensionabstractA fundamental trade-off between effectiveness and efficiency needs to be balanced when designing an online question answering system. Effectiveness comes from sophisticated functions such as extractive machine reading comprehension (MRC), while efficiency is obtained from improvements in preliminary retrieval components such as candidate document selection and paragraph ranking. Given the complexity of the real-world multi-document MRC scenario, it is difficult to jointly optimize both in an end-to-end system. To address this problem, we develop a novel deep cascade learning model, which progressively evolves from the documentlevel and paragraph-level ranking of candidate texts to more precise answer extraction with machine reading comprehension. Specifically, irrelevant documents and paragraphs are first filtered out with simple functions for efficiency consideration. Then we jointly train three modules on the remaining texts for better tracking the answer: the document extraction, the paragraph extraction and the answer extraction. Experiment results show that the proposed method outperforms the previous state-of-the-art methods on two large-scale multidocument benchmark datasets, i.e., TriviaQA and DuReader. In addition, our online system can stably serve typical scenarios with millions of daily requests in less than 50ms. Ming Yan 0008, Jiangnan Xia, Chen Wu 0006, Bin Bi, Zhongzhou Zhao, Ji Zhang 0011, Luo Si, Rui Wang 0005, Wei Wang 0225, Haiqing Chen |
AAAI | 1 |
| 2019 | Incorporating Relation Knowledge into Commonsense Reading Comprehension with Multi-task LearningabstractThis paper focuses on how to take advantage of external relational knowledge to improve machine reading comprehension (MRC) with multi-task learning. Most of the traditional methods in MRC assume that the knowledge used to get the correct answer generally exists in the given documents. However, in real-world task, part of knowledge may not be mentioned and machines should be equipped with the ability to leverage external knowledge. In this paper, we integrate relational knowledge into MRC model for commonsense reasoning. Specifically, based on a pre-trained language model (LM), We design two auxiliary relation-aware tasks to predict if there exists any commonsense relation and what is the relation type be-tween two words, in order to better model the interactions between document and candidate answer option. We conduct experiments on two multi-choice benchmark datasets: the SemEval-2018 Task11 and the Cloze Story Test. The experimental results demonstrate the effectiveness of the proposed method, which achieves superior performance compared with the comparable baselines on both datasets. Jiangnan Xia, Chen Wu 0006, Ming Yan 0008 |
CIKM | 3 |
| 2019 | Incorporating External Knowledge into Machine Reading for Generative Question AnsweringabstractBin Bi, Chen Wu, Ming Yan, Wei Wang, Jiangnan Xia, Chenliang Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bin Bi, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Jiangnan Xia, Chenliang Li 0003 |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Multi-Granularity Hierarchical Attention Fusion Networks for Reading Comprehension and Question AnsweringabstractThis paper describes a novel hierarchical attention network for reading comprehension style question answering, which aims to answer questions for a given narrative paragraph.In the proposed method, attention and fusion are conducted horizontally and vertically across layers at different levels of granularity between question and paragraph.Specifically, it first encode the question and paragraph with fine-grained language embeddings, to better capture the respective representations at semantic level.Then it proposes a multi-granularity fusion approach to fully fuse information from both global and attended representations.Finally, it introduces a hierarchical attention network to focuses on the answer span progressively with multi-level softalignment.Extensive experiments on the large-scale SQuAD and TriviaQA datasets validate the effectiveness of the proposed method.At the time of writing the paper (Jan.12th 2018), our model achieves the first position on the SQuAD leaderboard for both single and ensemble models.We also achieves state-of-the-art results on TriviaQA, AddSent and AddOne-Sent datasets. Wei Wang 0225, Chen Wu 0006, Ming Yan 0008 |
ACL (1) | 3 |
| 2018 | Understanding Dynamic Cross-OSN Associations for Cold-Start RecommendationabstractOnline social networks (OSNs) have become an essential part of people's daily life, and an increasing number of users are now using multiple OSNs for different social media services simultaneously. As a result, user's interests and preferences usually distribute in different OSNs. While most of the existing work mainly aggregates the distributed user behaviors or features directly, recently very few efforts have been focused on understanding the cross-OSN association from collective user behaviors. In this paper, we go one step further to consider the dynamic characteristic of user behaviors and propose a dynamic cross-OSN association mining framework. In this framework, dynamic user modeling is first conducted to capture the drift of user interest in each OSN. A session-based factorization method is then proposed to establish the cross-OSN association in a dynamic manner, by incrementally updating the derived association each time a new session of data arrives. Based on the derived dynamic association, we finally design a cold-start YouTube video recommendation application, by only utilizing users' behaviors in Twitter. Experiments are conducted using real-world user data from Twitter and YouTube. The results demonstrate the effectiveness of this proposed framework in capturing the underlying association between different OSNs and achieving superior cold-start recommendation performance. Jitao Sang 0001, Ming Yan 0008, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2017 | Session-aware Information Embedding for E-commerce Product RecommendationabstractMost of the existing recommender systems assume that user's visiting history can be constantly recorded. However, in recent online services, the user identification may be usually unknown and only limited online user behaviors can be used. It is of great importance to model the temporal online user behaviors and conduct recommendation for the anonymous users. In this paper, we propose a list-wise deep neural network based architecture to model the limited user behaviors within each session. To train the model efficiently, we first design a session embedding method to pre-train a session representation, which incorporates different kinds of user search behaviors such as clicks and views. Based on the learnt session representation, we further propose a list-wise ranking model to generate the recommendation result for each anonymous user session. We conduct quantitative experiments on a recently published dataset from an e-commerce company. The evaluation results validate the effectiveness of the proposed method, which can outperform the state-of-the-art. Chen Wu 0006, Ming Yan 0008 |
CIKM | 2 |
| 2016 | A Unified Video Recommendation by Cross-Network User ModelingabstractOnline video sharing sites are increasingly encouraging their users to connect to the social network venues such as Facebook and Twitter, with goals to boost user interaction and better disseminate the high-quality video content. This in turn provides huge possibilities to conduct cross-network collaboration for personalized video recommendation. However, very few efforts have been devoted to leveraging users’ social media profiles in the auxiliary network to capture and personalize their video preferences, so as to recommend videos of interest. In this article, we propose a unified YouTube video recommendation solution by transferring and integrating users’ rich social and content information in Twitter network. While general recommender systems often suffer from typical problems like cold-start and data sparsity, our proposed recommendation solution is able to effectively learn from users’ abundant auxiliary information on Twitter for enhanced user modeling and well address the typical problems in a unified framework. In this framework, two stages are mainly involved: (1) auxiliary-network data transfer, where user preferences are transferred from an auxiliary network by learning cross-network knowledge associations; and (2) cross-network data integration, where transferred user preferences are integrated with the observed behaviors on a target network in an adaptive fashion. Experimental results show that the proposed cross-network collaborative solution achieves superior performance not only in terms of accuracy, but also in improving the diversity and novelty of the recommended videos. Ming Yan 0008, Jitao Sang 0001, Changsheng Xu, M. Shamim Hossain |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2015 | Unified YouTube Video Recommendation via Cross-network CollaborationabstractThe ever growing number of videos on YouTube makes recommendation an important way to help users explore interesting videos. Similar to general recommender systems, YouTube video recommendation suffers from typical problems like new user, cold-start, data sparsity, etc. In this paper, we propose a unified YouTube video recommendation solution via cross-network collaboration: users' auxiliary information on Twitter are exploited to address the typical problems in single network-based recommendation solutions. The proposed two-stage solution first transfers user preferences from auxiliary network by learning cross-network behavior correlations, and then integrates the transferred preferences with the observed behaviors on target network in an adaptive fashion. Experimental results show that the proposed cross-network collaborative solution achieves superior performance not only in term of accuracy, but also in improving the diversity and novelty of the recommended videos. Ming Yan 0008, Jitao Sang 0001, Changsheng Xu |
ICMR | 1 |
| 2015 | YouTube Video Promotion by Cross-Network Association: @Britney to Advertise Gangnam StyleabstractThe emergence and rapid proliferation of various social media networks have reshaped the way how video contents are generated, distributed, and consumed in traditional video sharing portals. Nowadays, online videos can be accessed from far beyond the internal mechanisms of the video sharing portals, such as internal search and front page highlight. Recent studies have found that external referrers, such as external search engines and other social media websites, arise to be the new and important portals to lead users to online videos. In this paper, we introduce a novel cross-network collaborative application to help drive the online traffic for given videos in the traditional video portal YouTube by leveraging the high propagation efficiency of the popular Twitter followees. Since YouTube videos and Twitter followees distribute on heterogeneous spaces, we present a cross-network association-based solution framework. In this framework, we first represent YouTube videos and Twitter followees in the corresponding topic spaces separately by employing generative topic models. Then, the cross-network topic spaces are associated from both semantic-based and network-based perspectives through the collective intelligence of the observed overlapped users. Based on the derived cross-network association, we finally match the query YouTube videos and candidate Twitter followees in the same topic space with a unified ranking method. The experiments on a real-world large-scale dataset of more than 2.2 million YouTube videos and 31.8 million tweets from 38,540 YouTube users and 39,400 Twitter users demonstrate the effectiveness and superiority of our solution in which network-based and semantic-based association are integrated. Ming Yan 0008, Jitao Sang 0001, Changsheng Xu, M. Shamim Hossain |
IEEE Trans. Multim. | 1 |
| 2014 | Mining Cross-network Association for YouTube Video PromotionabstractWe introduce a novel cross-network collaborative problem in this work: given YouTube videos, to find optimal Twitter followees that can maximize the video promotion on Twitter. Since YouTube videos and Twitter followees distribute on heterogeneous spaces, we present a cross-network association-based solution framework. Three stages are addressed: (1) heterogeneous topic modeling, where YouTube videos and Twitter followees are modeled in topic level; (2) cross-network topic association, where the overlapped users are exploited to conduct cross-network topic distribution transfer; and (3) referrer identification, where the query YouTube video and candidate Twitter followees are matched in the same topic space. Different methods in each stage are designed and compared by qualitative as well as quantitative experiments. Based on the proposed framework, we also discuss the potential applications, extensions, and suggest some principles for future heterogeneous social media utilization and cross-network collaborative applications. Ming Yan 0008, Jitao Sang 0001, Changsheng Xu |
ACM Multimedia | 1 |
| 2014 | Twitter is Faster: Personalized Time-Aware Video Recommendation from Twitter to YouTubeabstractTraditional personalized video recommendation methods focus on utilizing user profile or user history behaviors to model user interests, which follows a static strategy and fails to capture the swift shift of the short-term interests of users. According to our cross-platform data analysis, the information emergence and propagation is faster in social textual stream-based platforms than that in multimedia sharing platforms at micro user level. Inspired by this, we propose a dynamic user modeling strategy to tackle personalized video recommendation issues in the multimedia sharing platform YouTube, by transferring knowledge from the social textual stream-based platform Twitter. In particular, the cross-platform video recommendation strategy is divided into two steps. (1) Real-time hot topic detection: the hot topics that users are currently following are extracted from users' tweets, which are utilized to obtain the related videos in YouTube. (2) Time-aware video recommendation: for the target user in YouTube, the obtained videos are ranked by considering the user profile in YouTube, time factor, and quality factor to generate the final recommendation list. In this way, the short-term (hot topics) and long-term (user profile) interests of users are jointly considered. Carefully designed experiments have demonstrated the advantages of the proposed method. Zhengyu Deng, Ming Yan 0008, Jitao Sang 0001, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Friend transfer: Cold-start friend recommendation with cross-platform transfer learning of social knowledgeabstractThe emergence of various and disparate social media platforms has opened opportunities for the research on cross-platform media analysis. This provides huge potentials to solve many challenging problems which cannot be well explored in one single platform. In this paper, we investigate into cross-platform social relation and behavior information to address the cold-start friend recommendation problem. In particular, we conduct an in-depth data analysis to examine what information can better transfer from one platform to another and the result demonstrates a strong correlation for the bidirectional relation and common contact behavior between our test platforms. Inspired by the observations, we design a random walk-based method to employ and integrate these convinced social information to boost friend recommendation performance. To validate the effectiveness of our cross-platform social transfer learning, we have collected a cross-platform dataset including 3,000 users with recognized accounts in both Flickr and Twitter. We demonstrate the effectiveness of the proposed friend transfer methods by promising results. Ming Yan 0008, Jitao Sang 0001, Tao Mei 0001, Changsheng Xu |
ICME | 1 |