Weiming Lu 0001

dblp:41/5305-1 · also Wei-ming Lu 0001 · DBLP profile ↗
← Back
103ranked-venue papers
13as first author
60since 2021 · last 2026
0000-0002-0200-9215ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 70 · 6 first-author · 51 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 3 since 2021Systems, architecture and hardware · 4 · 1 first-author
YearPublicationVenuePosition
2026 Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
abstract
Graphical User Interface (GUI) grounding, the task of mapping natural language instructions to precise screen coordinates, is fundamental to autonomous GUI agents. While existing methods achieve strong performance through extensive supervised training or reinforcement learning with labeled rewards, they remain constrained by the cost and availability of pixel-level annotations. We observe that when models generate multiple predictions for the same GUI element, the spatial overlap patterns reveal implicit confidence signals that can guide more accurate localization. Leveraging this insight, GUI-RC (Region Consistency), a test-time scaling method that constructs spatial voting grids from multiple sampled predictions to identify consensus regions where models show highest agreement. Without any training, GUI-RC improves accuracy by 2-3% across various architectures on ScreenSpot benchmarks. We further introduce GUI-RCPO (Region Consistency Policy Optimization), transforming these consistency patterns into rewards for test-time reinforcement learning. By computing how well each prediction aligns with the collective consensus, GUI-RCPO enables models to iteratively refine their outputs on unlabeled data during inference. Extensive experiments demonstrate the generality of our approach: using only 1,272 unlabeled data, GUI-RCPO achieves 3-6% accuracy improvements across various architectures on ScreenSpot benchmarks. Our approach reveals the untapped potential of test-time scaling and test-time reinforcement learning for GUI grounding, offering a promising path toward more data-efficient GUI agents.
Fei Tang 0005, Zhengxi Lu, Chang Zong, Weiming Lu 0001, Shengpei Jiang, Yongliang Shen 0001
AAAI6
2026 Reality vs Counterfactual: Multi-World Contrastive Reinforcement Learning for Enhancing MLLM's Theory of Mind in Egocentric Videos
abstract
Theory of Mind (ToM) refers to the ability to infer others' mental states, which is an essential capability for embodied AI agents to effectively collaborate and interact with humans. While improving Large Language Models' ability to reason about characters' mental states in text-based stories/dialogues has been extensively studied, enhancing Multimodal Large Language Models' ToM capabilities, particularly in egocentric video from an embodied perspective, remains unexplored. In this paper, we propose a contrastive Reinforcement Learning (RL) paradigm that explicitly encourages models to leverage temporal and causal evolutionary patterns in user action sequences to infer user's mental states (goals, beliefs, and potential next actions). Evaluation results on in-domain and out-of-domain demonstrate that our method achieves performance improvements of (+30.00%, +2.00%) and (+5.83%, +5.00%) compared to the backbone model and vanilla Group Relative Policy Optimization (GRPO) model, respectively. Additionally, we compare the performance of two post-training paradigms (Supervise Fine-Tuning and RL) and systematically analyze the reasoning trajectories across the base model, vanilla GRPO model, and our proposed method.
Guiyang Hou, Yihui Fu, Wenqi Zhang 0001, Yongliang Shen 0001, Weiming Lu 0001
AAAI8
2026 GUI-G²: Gaussian Reward Modeling for GUI Grounding
abstract
Graphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use binary rewards that treat elements as hit-or-miss targets, creating sparse signals that ignore the continuous nature of spatial interactions. Motivated by human clicking behavior that naturally forms Gaussian distributions centered on target elements, we introduce GUI Gaussian Grounding Rewards (GUI-G2), a principled reward framework that models GUI elements as continuous Gaussian distributions across the interface plane. GUI-G2 incorporates two synergistic mechanisms: Gaussian point rewards model precise localization through exponentially decaying distributions centered on element centroids, while coverage rewards assess spatial alignment by measuring the overlap between predicted Gaussian distributions and target regions. To handle diverse element scales, we develop an adaptive variance mechanism that calibrates reward distributions based on element dimensions. This framework transforms GUI grounding from sparse binary classification to dense continuous optimization, where Gaussian distributions generate rich gradient signals that guide models toward optimal interaction positions. Extensive experiments across ScreenSpot, ScreenSpot-v2, and ScreenSpot-Pro benchmarks demonstrate that GUI-G2, substantially outperforms state-of-the-art method UI-TARS-72B, with the most significant improvement of 24.7% on ScreenSpot-Pro. Our analysis reveals that continuous modeling provides superior robustness to interface variations and enhanced generalization to unseen layouts, establishing a new paradigm for spatial reasoning in GUI interaction tasks.
Fei Tang 0005, Zhangxuan Gu, Zhengxi Lu, Shuheng Shen, Changhua Meng, Wen Wang 0009, Wenqi Zhang 0001, Yongliang Shen 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
AAAI10
2026 Leveraging Outline-Optimized Generative Interactions and Critique for Self-Refining Outlines with Reinforcement Learning
abstract
Hengwei Liu, Haoyuan Ma, Qingqing Lyu, Daoxin Zhang, Yao Hu, Yongliang Shen, Yin Zhang, Weiming Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hengwei Liu, Qingqing Lyu, Daoxin Zhang, Yongliang Shen 0001, Yin Zhang 0006, Weiming Lu 0001
ACL (1)8
2026 UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization
abstract
Zhengxi Lu, Fei Tang, Guangyi Liu, Jin Ma, Kaitao Song, Xu Tan, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengxi Lu, Fei Tang 0005, Kaitao Song, Xu Tan 0003, Wenqi Zhang 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang, Yongliang Shen 0001
ACL (1)8
2026 Experience-driven Multi-turn Reinforcement Learning for GUI Agents
abstract
Zhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen, Haiyang Xu, Ziwei Zheng, Weiming Lu, Ming Yan, Fei Huang, Jun Xiao, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengxi Lu, Jiabo Ye, Fei Tang 0005, Yongliang Shen 0001, Haiyang Xu 0001, Ziwei Zheng, Weiming Lu 0001, Ming Yan 0008, Fei Huang 0002, Jun Xiao 0001, Yueting Zhuang
ACL (1)7
2026 AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMs
abstract
Qingqing Lyu, Linjuan Wu, Yongliang Shen, Hengwei Liu, Hao Li, Shengpei Jiang, Yin Zhang, Weiming Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qingqing Lyu, Linjuan Wu, Yongliang Shen 0001, Hengwei Liu, Shengpei Jiang, Yin Zhang 0006, Weiming Lu 0001
ACL (1)8
2026 CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution
abstract
Teng Pan, Yuchen Yan, Zixuan Wang, Ruiqing Zhang, Guiyang Hou, Wenqi Zhang, Weiming Lu, Jun Xiao, Yongliang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Teng Pan, Ruiqing Zhang, Guiyang Hou, Wenqi Zhang 0001, Weiming Lu 0001, Jun Xiao 0001, Yongliang Shen 0001
ACL (1)7
2026 Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
abstract
Haolei Xu, Haiwen Hong, Hongxing Li, Rui Zhou, Yang Zhang, Longtao Huang, Hui Xue, Yongliang Shen, Weiming Lu, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Haolei Xu, Haiwen Hong, Longtao Huang, Hui Xue 0001, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang
ACL (1)9
2026 Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese
abstract
Language is a vehicle for thought, intricately tied to sounds, symbols, and meaning.However, most large language model (LLM) research focuses on meaning (semantics) and symbols (spelling) while largely overlooking sounds.Existing benchmarks on LLMs' phonological abilities are either solvable through rote memorization or intertwined with other abilities, making them inadequate to measure LLMs' genuine ability in phonological understanding.Here, we present Phun-Bench, a purpose-built Chinese benchmark with diverse tasks and settings across three dimensions (Homophony, Rhyme, and Phonetic Similarity), designed to systematically evaluate LLMs' phonological understanding.Our results show that while LLMs excel at recalling correct pronunciations, they generally struggle to leverage phonological knowledge in the flexible and intuitive way that human speakers do.Moreover, through detailed analyses, we propose a hypothesis regarding the underlying mechanism of LLMs' phonological understanding and "perception", highlighting an underexplored frontier for future research.1
Xing Yue, Yongliang Shen 0001, Weiming Lu 0001
ACL (1)3
2026 Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
abstract
Wenqi Zhang, Mengna Wang, Gangao Liu, Huixin Xu, Yiwei Jiang, Yongliang Shen, Guiyang Hou, Zhe Zheng, Hang Zhang, Xin Li, Jiajun Liu, Weiming Lu, Peng Li, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wenqi Zhang 0001, Mengna Wang, Gangao Liu, Huixin Xu, Yongliang Shen 0001, Guiyang Hou, Xin Li 0056, Weiming Lu 0001, Peng Li 0031, Yueting Zhuang
ACL (1)12
2026 Structural-temporal coupling anomaly detection with dynamic graph transformer
Chang Zong, Yueting Zhuang, Jian Shao 0001, Weiming Lu 0001
Data Min. Knowl. Discov.4
2025 G2LDetect: A Global-to-Local Approach for Hallucination Detection
abstract
Hallucination detection has attracted considerable interest due to the tendency of language models to generate texts that contain hallucinations. Most existing methods start with specific local details directly extracted from text, then aggregate to form the final conclusion. However, this direct extraction approach ignores the global context, leading to isolated details, and is prone to missed or over-detections. In this paper, we present a global-to-local approach for hallucination detection (G2LDetect), which considers the global information of the text before identifying local details. We first construct a global representation of the text by transforming it into a hierarchical tree structure. Afterward, we obtain specific local details from the global tree representation using path-wise identification and perform detection on them. This global-to-local detection process ensures that local details are context-aware and complete, thus making more accurate and reliable detection results. Experimental results show that our global-to-local method outperforms existing methods, especially for longer texts.
Xiaoxia Cheng, Zeqi Tan, Weiming Lu 0001
AAAI4
2025 STaR-SQL: Self-Taught Reasoner for Text-to-SQL
abstract
Generating step-by-step "chain-of-thought" rationales has proven effective for improving the performance of large language models on complex reasoning tasks.However, applying such techniques to structured tasks, such as text-to-SQL, remains largely unexplored.In this paper, we introduce Self-Taught Reasoner for text-to-SQL (STaR-SQL), a novel approach that reframes SQL query generation as a reasoningdriven process.Our method prompts the LLM to produce detailed reasoning steps for SQL queries and fine-tunes it on rationales that lead to correct outcomes.Unlike traditional methods, STaR-SQL dedicates additional test-time computation to reasoning, thereby positioning LLMs as spontaneous reasoners rather than mere prompt-based agents.To further scale the inference process, we incorporate an outcomesupervised reward model (ORM) as a verifier, which enhances SQL query accuracy.Experimental results on the challenging Spider benchmark demonstrate that STaR-SQL significantly improves text-to-SQL performance, achieving an execution accuracy of 86.6%.This surpasses a few-shot baseline by 31.6% and a baseline fine-tuned to predict answers directly by 18.0%.Additionally, STaR-SQL outperforms agent-like prompting methods that leverage more powerful yet closed-source models such as GPT-4.These findings underscore the potential of reasoning-augmented training for structured tasks and open the door to extending self-improving reasoning models to text-to-SQL generation and beyond.
Mingqian He, Yongliang Shen 0001, Wenqi Zhang 0001, Qiuying Peng, Weiming Lu 0001
ACL (1)6
2025 From English to Second Language Mastery: Enhancing LLMs with Cross-Lingual Continued Instruction Tuning
abstract
Supervised Fine-Tuning (SFT) with translated instruction data effectively adapts Large Language Models (LLMs) from English to non-English languages.We introduce Cross-Lingual Continued Instruction Tuning (X-CIT), which fully leverages translation-based parallel instruction data to enhance cross-lingual adaptability.X-CIT emulates the human process of second language acquisition and is guided by Chomsky's Principles and Parameters Theory.It first fine-tunes the LLM on English instruction data to establish foundational capabilities (i.e.Principles), then continues with target language translation and customized chatinstruction data to adjust "parameters" specific to the target language.This chat-instruction data captures alignment information in translated parallel data, guiding the model to initially think and respond in its native language before transitioning to the target language.To further mimic human learning progression, we incorporate Self-Paced Learning (SPL) during continued training, allowing the model to advance from simple to complex tasks.Implemented on Llama-2-7B across five languages, X-CIT was evaluated against three objective benchmarks and an LLM-as-a-judge benchmark, improving the strongest baseline by an average of 1.97% and 8.2% in these two benchmarks, respectively.
Linjuan Wu, Baosong Yang, Weiming Lu 0001
ACL (1)4
2025 Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training
abstract
Large language models (LLMs) exhibit remarkable multilingual capabilities despite Englishdominated pre-training, attributed to crosslingual mechanisms during pre-training.Existing methods for enhancing cross-lingual transfer remain constrained by parallel resources, suffering from limited linguistic and domain coverage.We propose Cross-lingual In-context Pre-training (CrossIC-PT), a simple and scalable approach that enhances cross-lingual transfer by leveraging semantically related bilingual texts via simple next-word prediction.We construct CrossIC-PT samples by interleaving semantic-related bilingual Wikipedia documents into a single context window.To access window size constraints, we implement a systematic segmentation policy to split long bilingual document pairs into chunks while adjusting the sliding window mechanism to preserve contextual coherence.We further extend data availability through a semantic retrieval framework to construct CrossIC-PT samples from web-crawled corpus.Experimental results demonstrate that CrossIC-PT improves multilingual performance on three models (Llama-3.1-8B,Qwen2.5-7B, and Qwen2.5-1.5B)across six target languages, yielding performance gains of 3.79%, 3.99%, and 1.95%, respectively, with additional improvements after data augmentation.
Linjuan Wu, Baosong Yang, Fei Huang 0002, Weiming Lu 0001
EMNLP7
2025 AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
abstract
Xuan Zhang, Yongliang Shen, Zhe Zheng, Linjuan Wu, Wenqi Zhang, Yuchen Yan, Qiuying Peng, Jun Wang, Weiming Lu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yongliang Shen 0001, Linjuan Wu, Wenqi Zhang 0001, Qiuying Peng, Weiming Lu 0001
EMNLP9
2025 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
abstract
Compared to image-text pair data, interleaved corpora enable Vision-Language Models (VLMs) to understand the world more naturally like humans. However, such existing datasets are crawled from webpage, facing challenges like low knowledge density, loose image-text relations, and poor logical coherence between images. On the other hand, the internet hosts vast instructional videos (e.g., online geometry courses) that are widely used by humans to learn foundational subjects, yet these valuable resources remain underexplored in VLM training. In this paper, we introduce a high-quality \textbf{multimodal textbook} corpus with richer foundational knowledge for VLM pretraining. It collects over 2.5 years of instructional videos, totaling 22,000 class hours. We first use an LLM-proposed taxonomy to systematically gather instructional videos. Then we progressively extract and refine visual (keyframes), audio (ASR), and textual knowledge (OCR) from the videos, and organize as an image-text interleaved corpus based on temporal order. Compared to its counterparts, our video-centric textbook offers more coherent context, richer knowledge, and better image-text alignment. Experiments demonstrate its superb pretraining performance, particularly in knowledge- and reasoning-intensive tasks like ScienceQA and MathVista. Moreover, VLMs pre-trained on our textbook exhibit outstanding interleaved context awareness, leveraging visual and textual cues in their few-shot context for task solving. Our code are available at https://github.com/DAMO-NLP-SG/multimodal_textbook.
Wenqi Zhang 0001, Xin Li 0056, Jiashuo Sun, Yongliang Shen 0001, Weiming Lu 0001, Deli Zhao, Yueting Zhuang, Lidong Bing
ICCV6
2025 SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
abstract
Large Language Models (LLMs) and Multimodal LLMs have shown promising capabilities for SVG processing, yet existing benchmarks suffer from limited real-world coverage, lack of complexity stratification, and fragmented evaluation paradigms. We introduce SVGenius, a comprehensive benchmark comprising 2,377 queries across three progressive dimensions: understanding, editing, and generation. Built on real-world data from 24 application domains with systematic complexity stratification, SVGenius evaluates models through 8 task categories and 18 metrics. We assess 22 mainstream models spanning different scales, architectures, training paradigms, and accessibility levels. Our analysis reveals that while proprietary models significantly outperform open-source counterparts, all models exhibit systematic performance degradation with increasing complexity, indicating fundamental limitations in current approaches; however, reasoning-enhanced training proves more effective than pure scaling for overcoming these limitations, though style transfer remains the most challenging capability across all model types. SVGenius establishes the first systematic evaluation framework for SVG processing, providing crucial insights for developing more capable vector graphics models and advancing automated graphic design applications. Appendix and supplementary materials (including all data and code) are available at https://zju-real.github.io/SVGenius.
Haolei Xu, Fei Tang 0005, Linjuan Wu, Wenqi Zhang 0001, Guiyang Hou, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang
ACM Multimedia12
2025 Legal Judgment Prediction based on Knowledge-enhanced Multi-Task and Multi-Label Text Classification
abstract
Ang Li, Yiquan Wu, Ming Cai, Adam Jatowt, Xiang Zhou, Weiming Lu, Changlong Sun, Fei Wu, Kun Kuang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ang Li 0049, Yiquan Wu 0001, Adam Jatowt, Weiming Lu 0001, Changlong Sun, Fei Wu 0001, Kun Kuang 0001
NAACL (Long Papers)6
2025 Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
abstract
Large language models (LLMs) have achieved remarkable progress on mathematical tasks through Chain-of-Thought (CoT) reasoning. However, existing mathematical CoT datasets often suffer from **Thought Leaps** due to experts omitting intermediate steps, which negatively impacts model learning and generalization. We propose the CoT Thought Leap Bridge Task, which aims to automatically detect leaps and generate missing intermediate reasoning steps to restore the completeness and coherence of CoT. To facilitate this, we constructed a specialized training dataset called **ScaleQM+**, based on the structured ScaleQuestMath dataset, and trained **CoT-Bridge** to bridge thought leaps. Through comprehensive experiments on mathematical reasoning benchmarks, we demonstrate that models fine-tuned on bridged datasets consistently outperform those trained on original datasets, with improvements of up to +5.87\% on NuminaMath. Our approach effectively enhances distilled data (+3.02\%) and provides better starting points for reinforcement learning (+3.1\%), functioning as a plug-and-play module compatible with existing optimization techniques. Furthermore, CoT-Bridge demonstrates improved generalization to out-of-domain logical reasoning tasks, confirming that enhancing reasoning completeness yields broadly applicable benefits.
Haolei Xu, Yongliang Shen 0001, Wenqi Zhang 0001, Guiyang Hou, Shengpei Jiang, Kaitao Song, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
NeurIPS8
2025 Let LRMs Break Free from Overthinking via Self-Braking Tuning
abstract
Large reasoning models (LRMs), such as OpenAI o1 and DeepSeek-R1, have significantly enhanced their reasoning capabilities by generating longer chains of thought, demonstrating outstanding performance across a variety of tasks. However, this performance gain comes at the cost of a substantial increase in redundant reasoning during the generation process, leading to high computational overhead and exacerbating the issue of overthinking. Although numerous existing approaches aim to address the problem of overthinking, they often rely on external interventions. In this paper, we propose a novel framework, **Self-Braking Tuning**(SBT), which tackles overthinking from the perspective of allowing the model to regulate its own reasoning process, thus eliminating the reliance on external control mechanisms. We construct a set of overthinking identification metrics based on standard answers and design a systematic method to detect redundant reasoning. This method accurately identifies unnecessary steps within the reasoning trajectory and generates training signals for learning self-regulation behaviors. Building on this foundation, we develop a complete strategy for constructing data with adaptive reasoning lengths and introduce an innovative braking prompt mechanism that enables the model to naturally learn when to terminate reasoning at an appropriate point. Experiments across mathematical benchmarks (AIME, AMC, MATH500, GSM8K) demonstrate that our method reduces token consumption by up to 60\% while maintaining comparable accuracy to unconstrained models.
Yongliang Shen 0001, Haolei Xu, Wenqi Zhang 0001, Kaitao Song, Jian Shao 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
NeurIPS8
2025 RexUniNLU: Recursive Method With Explicit Schema Instructor for Universal Natural Language Understanding
abstract
Information Extraction (IE) and Text Classification (CLS) serve as the fundamental pillars of NLU, with both disciplines relying on analyzing input sequences to categorize outputs into pre-established schemas. However, there is no existing encoder-based model that can unify IE and CLS tasks from this perspective. To fully explore the foundation shared within NLU tasks, we have proposed arecursive method with explicit schema instructor for universal NLU. Specifically, we firstly redefine the true universal information extraction (UIE) with a formal formulation that covers almost all extraction schemas, including quadruples and quintuples which remain unsolved for previous UIE models. Then, we expands the formulation to all CLS and multi-modal NLU tasks. Based on that, we introduce RexUniNLU, an universal NLU solution that employs explicit schema constraints for IE and CLS, which encompasses all IE and CLS tasks and prevent incorrect connections between schema and input sequence. To avoid interference between different schemas, we reset the position ids and attention mask matrices. Extensive experiments are conducted on IE, CLS in both English and Chinese, and multi-modality, revealing the effectiveness and superiority. Our codes are publicly released athttps://modelscope.cn/models/iic/nlp_deberta_rex-uninlu_chinese-base.
Yangyang Kang, Fubang Zhao, Kun Kuang 0001, Weiming Lu 0001, Changlong Sun, Fei Wu 0001
IEEE Trans. Knowl. Data Eng.6
2024 De-biased Attention Supervision for Text Classification with Causality
abstract
In text classification models, while the unsupervised attention mechanism can enhance performance, it often produces attention distributions that are puzzling to humans, such as assigning high weight to seemingly insignificant conjunctions. Recently, numerous studies have explored Attention Supervision (AS) to guide the model toward more interpretable attention distributions. However, such AS can impact classification performance, especially in specialized domains. In this paper, we address this issue from a causality perspective. Firstly, we leverage the causal graph to reveal two biases in the AS: 1) Bias caused by the label distribution of the dataset. 2) Bias caused by the words' different occurrence ranges that some words can occur across labels while others only occur in a particular label. We then propose a novel De-biased Attention Supervision (DAS) method to eliminate these biases with causal techniques. Specifically, we adopt backdoor adjustment on the label-caused bias and reduce the word-caused bias by subtracting the direct causal effect of the word. Through extensive experiments on two professional text classification datasets (e.g., medicine and law), we demonstrate that our method achieves improved classification accuracy along with more coherent attention distributions.
Yiquan Wu 0001, Ziyu Zhao 0001, Weiming Lu 0001, Changlong Sun, Fei Wu 0001, Kun Kuang 0001
AAAI4
2024 Learning Global Controller in Latent Space for Parameter-Efficient Fine-Tuning
abstract
Zeqi Tan, Yongliang Shen, Xiaoxia Cheng, Chang Zong, Wenqi Zhang, Jian Shao, Weiming Lu, Yueting Zhuang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zeqi Tan, Yongliang Shen 0001, Xiaoxia Cheng, Chang Zong, Wenqi Zhang 0001, Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang
ACL (1)7
2024 Insert or Attach: Taxonomy Completion via Box Embedding
abstract
Taxonomy completion, enriching existing taxonomies by inserting new concepts as parents or attaching them as children, has gained significant interest.Previous approaches embed concepts as vectors in Euclidean space, which makes it difficult to model asymmetric relations in taxonomy.In addition, they introduce pseudo-leaves to convert attachment cases into insertion cases, leading to an incorrect bias in network learning dominated by numerous pseudo-leaves.Addressing these, our framework, TAXBOX, leverages box containment and center closeness to design two specialized geometric scorers within the box embedding space.These scorers are tailored for insertion and attachment operations and can effectively capture intrinsic relationships between concepts by optimizing on a granular box constraint loss.We employ a dynamic ranking loss mechanism to balance the scores from these scorers, allowing adaptive adjustments of insertion and attachment scores.Experiments on four real-world datasets show that TAXBOX significantly outperforms previous methods, yielding substantial improvements over prior methods in real-world datasets, with average performance boosts of 6.7%, 34.9%, and 51.4% in MRR, Hit@1, and Prec@1, respectively.
Yongliang Shen 0001, Wenqi Ren, Jietian Guo, Shiliang Pu, Weiming Lu 0001
ACL (1)6
2024 Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
abstract
Wenqi Zhang, Yongliang Shen, Linjuan Wu, Qiuying Peng, Jun Wang, Yueting Zhuang, Weiming Lu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenqi Zhang 0001, Yongliang Shen 0001, Linjuan Wu, Qiuying Peng, Yueting Zhuang, Weiming Lu 0001
ACL (1)7
2024 Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
abstract
Wenqi Zhang, Ke Tang, Hai Wu, Mengna Wang, Yongliang Shen, Guiyang Hou, Zeqi Tan, Peng Li, Yueting Zhuang, Weiming Lu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenqi Zhang 0001, Mengna Wang, Yongliang Shen 0001, Guiyang Hou, Zeqi Tan, Peng Li 0031, Yueting Zhuang, Weiming Lu 0001
ACL (1)10
2024 Advancing Process Verification for Large Language Models via Tree-Based Preference Learning
abstract
Large Language Models (LLMs) have demonstrated remarkable potential in handling complex reasoning tasks by generating step-by-step rationales.Some methods have proven effective in boosting accuracy by introducing extra verifiers to assess these paths.However, existing verifiers, typically trained on binarylabeled reasoning paths, fail to fully utilize the relative merits of intermediate steps, thereby limiting the effectiveness of the feedback provided.To overcome this limitation, we propose Tree-based Preference Learning Verifier (Tree-PLV), a novel approach that constructs reasoning trees via a best-first search algorithm and collects step-level paired data for preference training.Compared to traditional binary classification, step-level preferences more finely capture the nuances between reasoning steps, allowing for a more precise evaluation of the complete reasoning path.We empirically evaluate Tree-PLV across a range of arithmetic and commonsense reasoning tasks, where it significantly outperforms existing benchmarks.For instance, Tree-PLV achieved substantial performance gains over the Mistral-7B selfconsistency baseline on GSM8K (67.55% → 82.79%), MATH (17.00% → 26.80%), CSQA (68.14% → 72.97%), and StrategyQA (82.86% → 83.25%).Additionally, our study explores the appropriate granularity for applying preference learning, revealing that step-level guidance provides feedback that better aligns with the evaluation of the reasoning process.
Mingqian He, Yongliang Shen 0001, Wenqi Zhang 0001, Zeqi Tan, Weiming Lu 0001
EMNLP5
2024 Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
abstract
Wenqi Zhang, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu, Yueting Zhuang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wenqi Zhang 0001, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen 0001, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu 0001, Yueting Zhuang
EMNLP10
2024 Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
abstract
Recent progress with LLM-based agents has shown promising results across various tasks.However, their use in answering questions from knowledge bases remains largely unexplored.Implementing a KBQA system using traditional methods is challenging due to the shortage of task-specific training data and the complexity of creating task-focused model structures.In this paper, we present Triad, a unified framework that utilizes an LLM-based agent with multiple roles for KBQA tasks.The agent is assigned three roles to tackle different KBQA subtasks: agent as a generalist for mastering various subtasks, as a decision maker for the selection of candidates, and as an advisor for answering questions with knowledge.Our KBQA framework is executed in four phases, involving the collaboration of the agent's multiple roles.We evaluated the performance of our framework using three benchmark datasets, and the results show that our framework outperforms state-of-the-art systems on the LC-QuAD and YAGO-QA benchmarks, yielding F1 scores of 11.8% and 20.7%, respectively.
Chang Zong, Weiming Lu 0001, Jian Shao 0001, Heng Chang, Yueting Zhuang
EMNLP3
2024 Shapley Value Guided Extractive Text Summarization
abstract
Extractive summarization is a process that extracts salient sentences from the original document to create a summary. Typically, the process relies on sentence labels, which indicate whether a sentence should be included in the summary or not. Widely used sentence labels are constructed by greedily selecting sentences one by one from the original document. However, this greedy manner obtains sub-optimal sentence labels and can’t adequately distinguish the importance of individual sentences. Exploring new labeling schemes to construct superior and informative sentence labels is necessary. In this work, we take a cooperative game perspective on sentence label construction. We treat sentences as different players and label construction as a cooperative game between them. In such a perspective, we design an optimal and soft label, termed the Shapley Value of sentences (SVS) to reasonably discern the importance of individual sentences. To alleviate the computational complexity of SVS, we adopt a Monte Carlo sampling-based method to approximate it. Experimental results demonstrate that our proposed SVS label performs better on existing benchmarks across sentence-level and summary-level, in supervised and zero-shot settings.
Xiaoxia Cheng, Weiming Lu 0001
ICASSP2
2024 TaskBench: Benchmarking Large Language Models for Task Automation
abstract
In recent years, the remarkable progress of large language models (LLMs) has sparked interest in task automation, which involves decomposing complex tasks described by user instructions into sub-tasks and invoking external tools to execute them, playing a central role in autonomous agents. However, there is a lack of systematic and standardized benchmarks to promote the development of LLMs in task automation. To address this, we introduce TaskBench, a comprehensive framework to evaluate the capability of LLMs in task automation. Specifically, task automation can be divided into three critical stages: task decomposition, tool selection, and parameter prediction. To tackle the complexities inherent in these stages, we introduce the concept of Tool Graph to represent decomposed tasks and adopt a back-instruct method to generate high-quality user instructions. We propose TaskEval, a multi-faceted evaluation methodology that assesses LLM performance across these three stages. Our approach combines automated construction with rigorous human verification, ensuring high consistency with human evaluation. Experimental results demonstrate that TaskBench effectively reflects the capabilities of various LLMs in task automation. It provides insights into model performance across different task complexities and domains, pushing the boundaries of what current models can achieve. TaskBench offers a scalable, adaptable, and reliable benchmark for advancing LLM-based autonomous agents.
Yongliang Shen 0001, Kaitao Song, Xu Tan 0003, Wenqi Zhang 0001, Kan Ren, Weiming Lu 0001, Dongsheng Li 0002, Yueting Zhuang
NeurIPS7
2024 Information Re-Organization Improves Reasoning in Large Language Models
abstract
Improving the reasoning capabilities of large language models (LLMs) has attracted considerable interest. Recent approaches primarily focus on improving the reasoning process to yield a more precise final answer. However, in scenarios involving contextually aware reasoning, these methods neglect the importance of first identifying logical relationships from the context before proceeding with the reasoning. This oversight could lead to a superficial understanding and interaction with the context, potentially undermining the quality and reliability of the reasoning outcomes. In this paper, we propose an information re-organization (\textbf{InfoRE}) method before proceeding with the reasoning to enhance the reasoning ability of LLMs. Our re-organization method involves initially extracting logical relationships from the contextual content, such as documents or paragraphs, and subsequently pruning redundant content to minimize noise. Then, we utilize the re-organized information in the reasoning process. This enables LLMs to deeply understand the contextual content by clearly perceiving these logical relationships, while also ensuring high-quality responses by eliminating potential noise. To demonstrate the effectiveness of our approach in improving the reasoning ability, we conduct experiments using Llama2-70B, GPT-3.5, and GPT-4 on various contextually aware multi-hop reasoning tasks. Using only a zero-shot setting, our method achieves an average absolute improvement of 4\% across all tasks, highlighting its potential to improve the reasoning performance of LLMs.
Xiaoxia Cheng, Zeqi Tan, Weiming Lu 0001
NeurIPS4
2024 Specialized Mathematical Solving by a Step-By-Step Expression Chain Generation
abstract
Math Solving requires both semantic understanding and relation reasoning. Most current approaches treat it as a translation task from natural language to mathematical symbols, generating tokens one by one. However, token-level generation is usually vulnerable when confronted with diverse annotations and complex reasoning. We consider the equation is an ordered combination of multiple sub-expressions, and math reasoning should be performed on the sub-expression level rather than the token level. We treat sub-expression as a minimum generative unit and minimum reasoning node. At each step, candidate sub-expression nodes are generated in parallel, and the whole reasoning chain is deduced by combining multiple nodes in order. Besides, we can obtain multiple valid reasoning chains by sub-expression searching, further improving interpretability and precision. Experiments on multilingual datasets show our method significantly outperforms the baselines. Additionally, our approach is more stable and efficient when faced with the challenges of diverse annotation and complex reasoning with limited resources. Moreover, our approach can be seamlessly integrated with large language models (LLMs), enhancing LLM's mathematical reasoning capabilities at a minimal cost. Experiments show a synergistic collaboration between general-purpose LLM and our specialized model yields superior performance.
Wenqi Zhang 0001, Yongliang Shen 0001, Guiyang Hou, Kuangyi Wang, Weiming Lu 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 DiffusionNER: Boundary Diffusion for Named Entity Recognition
abstract
In this paper, we propose DIFFUSIONNER, which formulates the named entity recognition task as a boundary-denoising diffusion process and thus generates named entities from noisy spans.During training, DIFFUSIONNER gradually adds noises to the golden entity boundaries by a fixed forward diffusion process and learns a reverse diffusion process to recover the entity boundaries.In inference, DIFFU-SIONNER first randomly samples some noisy spans from a standard Gaussian distribution and then generates the named entities by denoising them with the learned reverse diffusion process.The proposed boundary-denoising diffusion process allows progressive refinement and dynamic sampling of entities, empowering DIFFUSIONNER with efficient and flexible entity generation capability.Experiments on multiple flat and nested NER datasets demonstrate that DIFFUSIONNER achieves comparable or even better performance than previous state-of-the-art models 1 .
Yongliang Shen 0001, Kaitao Song, Xu Tan 0003, Dongsheng Li 0002, Weiming Lu 0001, Yueting Zhuang
ACL (1)5
2023 PromptNER: Prompt Locating and Typing for Named Entity Recognition
abstract
Yongliang Shen, Zeqi Tan, Shuhui Wu, Wenqi Zhang, Rongsheng Zhang, Yadong Xi, Weiming Lu, Yueting Zhuang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yongliang Shen 0001, Zeqi Tan, Shuhui Wu, Wenqi Zhang 0001, Yadong Xi, Weiming Lu 0001, Yueting Zhuang
ACL (1)7
2023 Graph Neural Networks with Geometric Edge Fusion and Point Downsampling for Drug-Target Interaction Prediction
abstract
Accurate and efficient prediction of Drug-Target Interactions (DTIs) can potentially accelerate the drug discovery process. We propose a framework, namely GeoPD-DTI, that utilizes Geometric Edge Fusion to effectively integrate distance information from various sources to model drugs and target proteins, and leverages Point Downsampling to reduce the computational burden of modeling protein 3D structures without compromising prediction accuracy. Moreover, to alleviate the ambiguity in substructure modeling of drug 3D molecular graphs, we introduce molecular fingerprints as a supplement to the 3D molecular graphs. Our framework achieves state-of-the-art performance on two benchmark datasets. We also visualize the edge weights learned by our model, which demonstrates clear patterns of interactions. Our code is available at https://github.com/Hienyriux/GeoPD-DTI.
Weiming Lu 0001
BIBM3
2023 MProto: Multi-Prototype Network with Denoised Optimal Transport for Distantly Supervised Named Entity Recognition
abstract
Distantly supervised named entity recognition (DS-NER) aims to locate entity mentions and classify their types with only knowledge bases or gazetteers and unlabeled corpus.However, distant annotations are noisy and degrade the performance of NER models.In this paper, we propose a noise-robust prototype network named MProto for the DS-NER task.Different from previous prototype-based NER methods, MProto represents each entity type with multiple prototypes to characterize the intra-class variance among entity representations.To optimize the classifier, each token should be assigned an appropriate ground-truth prototype and we consider such token-prototype assignment as an optimal transport (OT) problem.Furthermore, to mitigate the noise from incomplete labeling, we propose a novel denoised optimal transport (DOT) algorithm.Specifically, we utilize the assignment result between Other class tokens and all prototypes to distinguish unlabeled entity tokens from true negatives.Experiments on several DS-NER benchmarks demonstrate that our MProto achieves state-of-the-art performance.The source code is now available on Github 1 .
Shuhui Wu, Yongliang Shen 0001, Zeqi Tan, Wenqi Ren, Jietian Guo, Shiliang Pu, Weiming Lu 0001
EMNLP7
2023 Struct-XLM: A Structure Discovery Multilingual Language Model for Enhancing Cross-lingual Transfer through Reinforcement Learning
abstract
Cross-lingual transfer learning heavily relies on well-aligned cross-lingual representations.The syntactic structure is recognized as beneficial for cross-lingual transfer, but limited researches utilize it for aligning representation in multilingual pre-trained language models (PLMs).Additionally, existing methods require syntactic labels that are difficult to obtain and of poor quality for low-resource languages.To address this gap, we propose Struct-XLM, a novel multilingual language model that leverages reinforcement learning (RL) to autonomously discover universal syntactic structures for improving the cross-lingual representation alignment of PLM.Struct-XLM integrates a policy network (PNet) and a translation ranking task.The PNet is designed to discover structural information and integrate it into the last layer of the PLM through the structural multi-head attention module to obtain structural representation.The translation ranking task obtains a delayed reward based on the structural representation to optimize the PNet while improving the alignment of cross-lingual representation.Experiments show the effectiveness of the proposed approach for enhancing cross-lingual transfer of multilingual PLM on the XTREME benchmark 1 .
Linjuan Wu, Weiming Lu 0001
EMNLP2
2023 Precedent-Enhanced Legal Judgment Prediction with LLM and Domain-Model Collaboration
abstract
Yiquan Wu, Siying Zhou, Yifei Liu, Weiming Lu, Xiaozhong Liu, Yating Zhang, Changlong Sun, Fei Wu, Kun Kuang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Yiquan Wu 0001, Siying Zhou, Weiming Lu 0001, Xiaozhong Liu 0001, Changlong Sun, Fei Wu 0001, Kun Kuang 0001
EMNLP4
2023 An Expression Tree Decoding Strategy for Mathematical Equation Generation
abstract
Generating mathematical equations from natural language requires an accurate understanding of the relations among math expressions.Existing approaches can be broadly categorized into token-level and expression-level generation.The former treats equations as a mathematical language, sequentially generating math tokens.Expression-level methods generate each expression one by one.However, each expression represents a solving step, and there naturally exist parallel or dependent relations between these steps, which are ignored by current sequential methods.Therefore, we integrate tree structure into the expression-level generation and advocate an expression tree decoding strategy.To generate a tree with expression as its node, we employ a layer-wise parallel decoding strategy: we decode multiple independent expressions (leaf nodes) in parallel at each layer and repeat parallel decoding layer by layer to sequentially generate these parent node expressions that depend on others.Besides, a bipartite matching algorithm is adopted to align multiple predictions with annotations for each layer.Experiments show our method outperforms other baselines, especially for these equations with complex structures.
Wenqi Zhang 0001, Yongliang Shen 0001, Qingpeng Nong, Zeqi Tan, Yanna Ma, Weiming Lu 0001
EMNLP6
2023 HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
abstract
Solving complicated AI tasks with different domains and modalities is a key step toward artificial general intelligence. While there are numerous AI models available for various domains and modalities, they cannot handle complicated AI tasks autonomously. Considering large language models (LLMs) have exhibited exceptional abilities in language understanding, generation, interaction, and reasoning, we advocate that LLMs could act as a controller to manage existing AI models to solve complicated AI tasks, with language serving as a generic interface to empower this. Based on this philosophy, we present HuggingGPT, an LLM-powered agent that leverages LLMs (e.g., ChatGPT) to connect various AI models in machine learning communities (e.g., Hugging Face) to solve AI tasks. Specifically, we use ChatGPT to conduct task planning when receiving a user request, select models according to their function descriptions available in Hugging Face, execute each subtask with the selected AI model, and summarize the response according to the execution results. By leveraging the strong language capability of ChatGPT and abundant AI models in Hugging Face, HuggingGPT can tackle a wide range of sophisticated AI tasks spanning different modalities and domains and achieve impressive results in language, vision, speech, and other challenging tasks, which paves a new way towards the realization of artificial general intelligence.
Yongliang Shen 0001, Kaitao Song, Xu Tan 0003, Dongsheng Li 0002, Weiming Lu 0001, Yueting Zhuang
NeurIPS5
2023 ML-LJP: Multi-Law Aware Legal Judgment Prediction
abstract
Legal judgment prediction (LJP) is a significant task in legal intelligence, which aims to assist the judges and determine the judgment result based on the case's fact description. The judgment result consists of law articles, charge, and prison term. The law articles serve as the basis for the charge and the prison term, which can be divided into two types, named as charge-related law article and term-related law article, respectively. Recently, many methods have been proposed and made tremendous progress in LJP. However, the existing methods only focus on the prediction of the charge-related law articles, ignoring the term-related law articles (e.g., laws about lenient treatment), which limits the performance in the prison term prediction. In this paper, following the actual legal process, we expand the law article prediction as a multi-label classification task that includes both the charge-related law articles and term-related law articles and propose a novel multi-law aware LJP (ML-LJP) method to improve the performance of LJP. Given the case's fact description, firstly, the label (e.g., law article and charge) definitions in the Code of Law are used to transform the representation of the fact into several label-specific representations and make the prediction of the law articles and the charge. To distinguish the similar content of different label definitions, contrastive learning is conducted in the training. Then, a graph attention network (GAT) is applied to learn the interactions among the multiple law articles for the prediction of the prison term. Since numbers (e.g., amount of theft and weight of drugs) are important for LJP but often ignored by conventional encoders, we design a corresponding number representation method to locate and better represent these effective numbers. Extensive experiments on real-world dataset show that our method achieves the best results compared to the state-of-the-art models, especially in the task of prison term prediction where ML-LJP achieves a 10.07% relative improvement over the best baseline.
Yiquan Wu 0001, Changlong Sun, Weiming Lu 0001, Fei Wu 0001, Kun Kuang 0001
SIGIR5
2023 Multi-Label Classification With Dual Tail-Node Augmentation for Drug Repositioning
abstract
Due to the lengthy and costly process of new drug discovery, increasing attention has been paid to drug repositioning, i.e., identifying new drug-disease associations. Current machine learning methods for drug repositioning mainly leverage matrix factorization or graph neural networks, and have achieved impressive performance. However, they often suffer from insufficient training labels of inter-domain associations, while ignore the intra-domain associations. Moreover, they often neglect the importance of tail nodes that have few known associations, which limits their effectiveness in drug repositioning. In this paper, we propose a novel multi-label classification model with dual Tail-Node Augmentation for Drug Repositioning (TNA-DR). We incorporate disease-disease similarity and drug-drug similarity information into k-nearest neighbor ( kNN) augmentation module and contrastive augmentation module, respectively, which effectively complements the weak supervision of drug-disease associations. Furthermore, before employing the two augmentation modules, we filter the nodes by their degrees, so that the two modules are only applied to tail nodes. We conduct 10-fold cross validation experiments on four different real-world datasets, and our model achieves the state-of-the-art performance on all the four datasets. We also demonstrate our model's capability of identifying drug candidates for new diseases and discovering potential new links between existing drugs and diseases.
Weiming Lu 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2023 Stable Prediction With Leveraging Seed Variable
abstract
In this paper, we focus on the problem of stable prediction across unknown test data, where the test distribution might be different from the training one and is always agnostic when model training. In such a case, previous machine learning methods might exploit subtly spurious correlations induced by non-causal variables in training data for prediction. Those spurious correlations are changeable across data, leading to instability of prediction across unknown test data. To address this problem, we propose a conditional independence test based algorithm to screen out part of non-causal features and reduce those spurious correlations for a more stable prediction by leveraging a seed variable. We show, both theoretically and with empirical experiments, that our algorithm can precisely screen out the isolated non-causal variables, which have no causal relationship with other variables, and remove the spurious correlations induced by them, increasing the stability of prediction across unknown test data. Extensive experiments on both synthetic and real-world datasets demonstrate that our algorithm outperforms state-of-the-art methods for stable prediction across unknown test data.
Kun Kuang 0001, Haotian Wang 0001, Ruoxuan Xiong, Runze Wu 0001, Weiming Lu 0001, Yueting Zhuang, Fei Wu 0001, Peng Cui 0001, Bo Li 0064
IEEE Trans. Knowl. Data Eng.6
2022 Parallel Instance Query Network for Named Entity Recognition
abstract
Yongliang Shen, Xiaobin Wang, Zeqi Tan, Guangwei Xu, Pengjun Xie, Fei Huang, Weiming Lu, Yueting Zhuang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yongliang Shen 0001, Xiaobin Wang, Zeqi Tan, Pengjun Xie, Fei Huang 0002, Weiming Lu 0001, Yueting Zhuang
ACL (1)7
2022 De-Bias for Generative Extraction in Unified NER Task
abstract
Named entity recognition (NER) is a fundamental task to recognize specific types of entities from a given sentence.Depending on how the entities appear in the sentence, it can be divided into three subtasks, namely, Flat NER, Nested NER, and Discontinuous NER.Among the existing approaches, only the generative model can be uniformly adapted to these three subtasks.However, when the generative model is applied to NER, its optimization objective is not consistent with the task, which makes the model vulnerable to the incorrect biases.In this paper, we analyze the incorrect biases in the generation process from a causality perspective and attribute them to two confounders: pre-context confounder and entityorder confounder.Furthermore, we design Intra-and Inter-entity Deconfounding Data Augmentation methods to eliminate the above confounders according to the theory of backdoor adjustment.Experiments show that our method can improve the performance of the generative NER model in various datasets.
Yongliang Shen 0001, Zeqi Tan, Yiquan Wu 0001, Weiming Lu 0001
ACL (1)5
2022 Molecular Substructure-Aware Network for Drug-Drug Interaction Prediction
abstract
Concomitant administration of drugs can cause drug-drug interactions (DDIs). Some drug combinations are beneficial, but other ones may cause negative effects which are previously unrecorded. Previous works on DDI prediction usually rely on hand-engineered domain knowledge, which is laborious to obtain. In this work, we propose a novel model, Molecular Substructure-Aware Network (MSAN), to effectively predict potential DDIs from molecular structures of drug pairs. We adopt a Transformer-like substructure extraction module to acquire a fixed number of representative vectors that are associated with various substructure patterns of the drug molecule. Then, interaction strength between the two drugs' substructures will be captured by a similarity-based interaction module. We also perform a substructure dropping augmentation before graph encoding to alleviate overfitting. Experimental results from a real-world dataset reveal that our proposed model achieves the state-of-the-art performance. We also show that the predictions of our model are highly interpretable through a case study.
Yongliang Shen 0001, Weiming Lu 0001
CIKM3
2022 Query-based Instance Discrimination Network for Relational Triple Extraction
abstract
Joint entity and relation extraction has been a core task in the field of information extraction.Recent approaches usually consider the extraction of relational triples from a stereoscopic perspective, either learning a relation-specific tagger or separate classifiers for each relation type.However, they still suffer from error propagation, relation redundancy and lack of highlevel connections between triples.To address these issues, we propose a novel query-based approach to construct instance-level representations for relational triples.By metric-based comparison between query embeddings and token embeddings, we can extract all types of triples in one step, thus eliminating the error propagation problem.In addition, we learn the instance-level representation of relational triples via contrastive learning.In this way, relational triples can not only enclose rich classlevel semantics but also access to high-order global connections.Experimental results show that our proposed method achieves the state of the art on five widely used benchmarks.
Zeqi Tan, Yongliang Shen 0001, Xuming Hu, Wenqi Zhang 0001, Xiaoxia Cheng, Weiming Lu 0001, Yueting Zhuang
EMNLP6
2022 Improving Complex Knowledge Base Question Answering via Question-to-Action and Question-to-Question Alignment
abstract
Complex knowledge base question answering can be achieved by converting questions into sequences of predefined actions.However, there is a significant semantic and structural gap between natural language and action sequences, which makes this conversion difficult.In this paper, we introduce an alignment-enhanced complex question answering framework, called ALCQA, which mitigates this gap through question-to-action alignment and question-toquestion alignment.We train a question rewriting model to align the question and each action, and utilize a pretrained language model to implicitly align the question and KG artifacts.Moreover, considering that similar questions correspond to similar action sequences, we retrieve top-k similar question-answer pairs at the inference stage through question-to-question alignment and propose a novel reward-guided action sequence selection strategy to select from candidate action sequences.We conduct experiments on CQA and WQSP datasets, and the results show that our approach outperforms state-of-the-art methods and obtains a 9.88% improvements in the F1 metric on CQA dataset.Our source code is available at https://github.com/TTTTTTTTy/ALCQA.
Yechun Tang, Xiaoxia Cheng, Weiming Lu 0001
EMNLP3
2022 Towards Interactivity and Interpretability: A Rationale-based Legal Judgment Prediction Framework
abstract
Legal judgment prediction (LJP) is a fundamental task in legal AI, which aims to assist the judge to hear the case and determine the judgment.The legal judgment usually consists of the law article, charge, and term of penalty.In the real trial scenario, the judge usually makes the decision step-by-step: first concludes the rationale according to the case's facts and then determines the judgment.Recently, many models have been proposed and made tremendous progress in LJP, but most of them adopt an endto-end manner that cannot be manually intervened by the judge for practical use.Moreover, existing models lack interpretability due to the neglect of rationale in the prediction process.Following the judge's real trial logic, in this paper, we propose a novel Rationale-based Legal Judgment Prediction (RLJP) framework.In the RLJP framework, the LJP process is split into two steps.In the first phase, the model generates the rationales according to the fact description.Then it predicts the judgment based on the fact and the generated rationales.Extensive experiments on a real-world dataset show RLJP achieves the best results compared to the stateof-the-art models.Meanwhile, the proposed framework provides good interactivity and interpretability which enables practical use. Fact DescriptionAfter the hearing, the court held the facts as follows: on December 3, 2014, the defendant A went to victim B's house and said he wanted to borrow B's money.But B refused, so A pushed B to the ground, bound B with a red soft cloth belt, and robbed B's cash of 1100.After the case, the parents of the defendant A returned 1100 yuan to B. On December 7, 2014, the defendant A surrendered to the court. Rationale Charge RationaleFor the purpose of illegal possession, the defendant A forcibly robbed B's property by means of violence. Penalty RationaleAfter the case, the defendant A voluntarily surrendered and actively returned the stolen goods, and may be given a lighter punishment as appropriate.
Yiquan Wu 0001, Weiming Lu 0001, Changlong Sun, Fei Wu 0001, Kun Kuang 0001
EMNLP3
2022 Prompting to Distill: Boosting Data-Free Knowledge Distillation via Reinforced Prompt
abstract
Data-free knowledge distillation (DFKD) conducts knowledge distillation via eliminating the dependence of original training data, and has recently achieved impressive results in accelerating pre-trained language models. At the heart of DFKD is to reconstruct a synthetic dataset by inverting the parameters of the uncompressed model. Prior DFKD approaches, however, have largely relied on hand-crafted priors of the target data distribution for the reconstruction, which can be inevitably biased and often incompetent to capture the intrinsic distributions. To address this problem, we propose a prompt-based method, termed as PromptDFD, that allows us to take advantage of learned language priors, which effectively harmonizes the synthetic sentences to be semantically and grammatically correct. Specifically, PromptDFD leverages a pre-trained generative model to provide language priors and introduces a reinforced topic prompter to control data synthesis, making the generated samples thematically relevant and semantically plausible, and thus friendly to downstream tasks. As shown in our experiments, the proposed method substantially improves the synthesis quality and achieves considerable improvements on distillation performance. In some cases, PromptDFD even gives rise to results on par with those from the data-driven knowledge distillation with access to the original training data.
Xinyin Ma, Xinchao Wang, Gongfan Fang, Yongliang Shen 0001, Weiming Lu 0001
IJCAI5
2022 Propose-and-Refine: A Two-Stage Set Prediction Network for Nested Named Entity Recognition
abstract
Nested named entity recognition (nested NER) is a fundamental task in natural language processing. Various span-based methods have been proposed to detect nested entities with span representations. However, span-based methods do not consider the relationship between a span and other entities or phrases, which is helpful in the NER task. Besides, span-based methods have trouble predicting long entities due to limited span enumeration length. To mitigate these issues, we present the Propose-and-Refine Network (PnRNet), a two-stage set prediction network for nested NER. In the propose stage, we use a span-based predictor to generate some coarse entity predictions as entity proposals. In the refine stage, proposals interact with each other, and richer contextual information is incorporated into the proposal representations. The refined proposal representations are used to re-predict entity boundaries and classes. In this way, errors in coarse proposals can be eliminated, and the boundary prediction is no longer constrained by the span enumeration length limitation. Additionally, we build multi-scale sentence representations, which better model the hierarchical structure of sentences and provide richer contextual information than token-level representations. Experiments show that PnRNet achieves state-of-the-art performance on four nested NER datasets and one flat NER dataset.
Shuhui Wu, Yongliang Shen 0001, Zeqi Tan, Weiming Lu 0001
IJCAI4
2022 A Closed-Loop Perception, Decision-Making and Reasoning Mechanism for Human-Like Navigation
abstract
Reliable navigation systems have a wide range of applications in robotics and autonomous driving. Current approaches employ an open-loop process that converts sensor inputs directly into actions. However, these open-loop schemes are challenging to handle complex and dynamic real-world scenarios due to their poor generalization. Imitating human navigation, we add a reasoning process to convert actions back to internal latent states, forming a two-stage closed loop of perception, decision-making, and reasoning. Firstly, VAE-Enhanced Demonstration Learning endows the model with the understanding of basic navigation rules. Then, two dual processes in RL-Enhanced Interaction Learning generate reward feedback for each other and collectively enhance obstacle avoidance capability. The reasoning model can substantially promote generalization and robustness, and facilitate the deployment of the algorithm to real-world robots without elaborate transfers. Experiments show our method is more adaptable to novel scenarios compared with state-of-the-art approaches.
Wenqi Zhang 0001, Peng Li 0031, Yongliang Shen 0001, Yanna Ma, Weiming Lu 0001
IJCAI8
2021 Locate and Label: A Two-stage Identifier for Nested Named Entity Recognition
abstract
Yongliang Shen, Xinyin Ma, Zeqi Tan, Shuai Zhang, Wen Wang, Weiming Lu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yongliang Shen 0001, Xinyin Ma, Zeqi Tan, Wen Wang 0009, Weiming Lu 0001
ACL/IJCNLP (1)6
2021 MuVER: Improving First-Stage Entity Retrieval with Multi-View Entity Representations
abstract
Entity retrieval, which aims at disambiguating mentions to canonical entities from massive KBs, is essential for many tasks in natural language processing.Recent progress in entity retrieval shows that the dual-encoder structure is a powerful and efficient framework to nominate candidates if entities are only identified by descriptions.However, they ignore the property that meanings of entity mentions diverge in different contexts and are related to various portions of descriptions, which are treated equally in previous works.In this work, we propose Multi-View Entity Representations (MuVER), a novel approach for entity retrieval that constructs multi-view representations for entity descriptions and approximates the optimal view for mentions via a heuristic searching method.Our method achieves the state-ofthe-art performance on ZESHEL and improves the quality of candidates on three standard Entity Linking datasets 1 .
Xinyin Ma, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Weiming Lu 0001
EMNLP (1)7
2021 A Sequence-to-Set Network for Nested Named Entity Recognition
abstract
Named entity recognition (NER) is a widely studied task in natural language processing. Recently, a growing number of studies have focused on the nested NER. The span-based methods, considering the entity recognition as a span classification task, can deal with nested entities naturally. But they suffer from the huge search space and the lack of interactions between entities. To address these issues, we propose a novel sequence-to-set neural network for nested NER. Instead of specifying candidate spans in advance, we provide a fixed set of learnable vectors to learn the patterns of the valuable spans. We utilize a non-autoregressive decoder to predict the final set of entities in one pass, in which we are able to capture dependencies between entities. Compared with the sequence-to-sequence method, our model is more suitable for such unordered recognition task as it is insensitive to the label order. In addition, we utilize the loss function based on bipartite matching to compute the overall training loss. Experimental results show that our proposed model achieves state-of-the-art on three nested NER corpora: ACE 2004, ACE 2005 and KBP 2017. The code is available at https://github.com/zqtan1024/sequence-to-set.
Zeqi Tan, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang
IJCAI4
2021 Heterogeneous Graph Neural Networks for Concept Prerequisite Relation Learning in Educational Data
abstract
Chenghao Jia, Yongliang Shen, Yechun Tang, Lu Sun, Weiming Lu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Chenghao Jia, Yongliang Shen 0001, Yechun Tang, Weiming Lu 0001
NAACL-HLT5
2021 A Trigger-Sense Memory Flow Framework for Joint Entity and Relation Extraction
abstract
Joint entity and relation extraction framework constructs a unified model to perform entity recognition and relation extraction simultaneously, which can exploit the dependency between the two tasks to mitigate the error propagation problem suffered by the pipeline model. Current efforts on joint entity and relation extraction focus on enhancing the interaction between entity recognition and relation extraction through parameter sharing, joint decoding, or other ad-hoc tricks (e.g., modeled as a semi-Markov decision process, cast as a multi-round reading comprehension task). However, there are still two issues on the table. First, the interaction utilized by most methods is still weak and uni-directional, which is unable to model the mutual dependency between the two tasks. Second, relation triggers are ignored by most methods, which can help explain why humans would extract a relation in the sentence. They’re essential for relation extraction but overlooked. To this end, we present a Trigger-Sense Memory Flow Framework (TriMF) for joint entity and relation extraction. We build a memory module to remember category representations learned in entity recognition and relation extraction tasks. And based on it, we design a multi-level memory flow attention mechanism to enhance the bi-directional interaction between entity recognition and relation extraction. Moreover, without any human annotations, our model can enhance relation trigger information in a sentence through a trigger sensor module, which improves the model performance and makes model predictions with better interpretation. Experiment results show that our proposed framework achieves state-of-the-art results by improves the relation F1 to 52.44% (+3.2%) on SciERC, 66.49% (+4.9%) on ACE05, 72.35% (+0.6%) on CoNLL04 and 80.66% (+2.3%) on ADE.
Yongliang Shen 0001, Xinyin Ma, Yechun Tang, Weiming Lu 0001
WWW4
2020 Adversarial Self-Supervised Data-Free Distillation for Text Classification
abstract
Large pre-trained transformer-based language models have achieved impressive results on a wide range of NLP tasks.In the past few years, Knowledge Distillation(KD) has become a popular paradigm to compress a computationally expensive model to a resource-efficient lightweight model.However, most KD algorithms, especially in NLP, rely on the accessibility of the original training dataset, which may be unavailable due to privacy issues.To tackle this problem, we propose a novel twostage data-free distillation method, named Adversarial self-Supervised Data-Free Distillation (AS-DFD), which is designed for compressing large-scale transformer-based models (e.g., BERT).To avoid text generation in discrete space, we introduce a Plug & Play Embedding Guessing method to craft pseudo embeddings from the teacher's hidden knowledge.Meanwhile, with a self-supervised module to quantify the student's ability, we adapt the difficulty of pseudo embeddings in an adversarial training manner.To the best of our knowledge, our framework is the first data-free distillation framework designed for NLP tasks.We verify the effectiveness of our method on several text classification datasets.
Xinyin Ma, Yongliang Shen 0001, Gongfan Fang, Chenghao Jia, Weiming Lu 0001
EMNLP (1)6
2020 Multi-hop Reading Comprehension across Documents with Path-based Graph Convolutional Network
abstract
Multi-hop reading comprehension across multiple documents attracts much attentions recently. In this paper, we propose a novel approach to tackle this multi-hop reading comprehension problem. Inspired by the human reasoning processing, we introduce a path-based graph with reasoning paths which extracted from supporting documents. The path-based graph can combine both the idea of the graph-based and path-based approaches, so it is better for multi-hop reasoning. Meanwhile, we propose Gated-GCN to accumulate evidences on the path-based graph, which contains a new question-aware gating mechanism to regulate the usefulness of information propagating across documents and add question information during reasoning. We evaluate our approach on WikiHop dataset, and our approach achieves the the-state-of-art accuracy against previous published approaches. Especially, our ensemble model surpasses the human performance by 4.2%.
Zeyun Tang, Yongliang Shen 0001, Xinyin Ma, Jiale Yu, Weiming Lu 0001
IJCAI6
2020 Boosting Cross-lingual Entity Alignment with Textual Embedding
Chenghao Jia, Yongliang Shen 0001, Xinyin Ma, Weiming Lu 0001
NLPCC (2)6
2020 Enrich cross-lingual entity links for online wikis via multi-modal semantic matching
Weiming Lu 0001, Xinyin Ma
Inf. Process. Manag.1
2020 EncyCatalogRec: catalog recommendation for encyclopedia article completion
abstract
Online encyclopedias such as Wikipedia provide a large and growing number of articles on many topics. However, the content of many articles is still far from complete. In this paper, we propose EncyCatalogRec, a system to help generate a more comprehensive article by recommending catalogs. First, we represent articles and catalog items as embedding vectors, and obtain similar articles via the locality sensitive hashing technology, where the items of these articles are considered as the candidate items. Then a relation graph is built from the articles and the candidate items. This is further transformed into a product graph. So, the recommendation problem is changed to a transductive learning problem in the product graph. Finally, the recommended items are sorted by the learning-to-rank technology. Experimental results demonstrate that our approach achieves state-of-the-art performance on catalog recommendation in both warm- and cold-start scenarios. We have validated our approach by a case study.
Weiming Lu 0001, Baogang Wei
Frontiers Inf. Technol. Electron. Eng.1
2019 Concept Extraction and Prerequisite Relation Learning from Educational Data
abstract
Prerequisite relations among concepts are crucial for educational applications. However, it is difficult to automatically extract domain-specific concepts and learn the prerequisite relations among them without labeled data.In this paper, we first extract high-quality phrases from a set of educational data, and identify the domain-specific concepts by a graph based ranking method. Then, we propose an iterative prerequisite relation learning framework, called iPRL, which combines a learning based model and recovery based model to leverage both concept pair features and dependencies among learning materials. In experiments, we evaluated our approach on two real-world datasets Textbook Dataset and MOOC Dataset, and validated that our approach can achieve better performance than existing methods. Finally, we also illustrate some examples of our approach.
Weiming Lu 0001, Jiale Yu, Chenhao Jia
AAAI1
2019 Incorporating External Knowledge to Boost Machine Comprehension Based Question Answering
Weiming Lu 0001, Zeyun Tang
ECIR (1)2
2019 Metro maps for efficient knowledge learning by summarizing massive electronic textbooks
Weiming Lu 0001, Pengkun Ma, Jiale Yu, Baogang Wei
Int. J. Document Anal. Recognit.1
2018 Active instance matching with pairwise constraints and its application to Chinese knowledge base construction
Weiming Lu 0001, Zhenyu Zhang 0008, Yueting Zhuang
Knowl. Inf. Syst.1
2018 Identifying Objective and Subjective Words via Topic Modeling
abstract
It is observed that distinct words in a given document have either strong or weak ability in delivering facts (i.e., the objective sense) or expressing opinions (i.e., the subjective sense) depending on the topics they associate with. Motivated by the intuitive assumption that different words have varying degree of discriminative power in delivering the objective sense or the subjective sense with respect to their assigned topics, a model named as dentified bjective- ubjective latent Dirichlet allocation (LDA) ( osLDA) is proposed in this paper. In the osLDA model, the simple Pólya urn model adopted in traditional topic models is modified by incorporating it with a probabilistic generative process, in which the novel "Bag-of-Discriminative-Words" (BoDW) representation for the documents is obtained; each document has two different BoDW representations with regard to objective and subjective senses, respectively, which are employed in the joint objective and subjective classification instead of the traditional Bag-of-Topics representation. The experiments reported on documents and images demonstrate that: 1) the BoDW representation is more predictive than the traditional ones; 2) osLDA boosts the performance of topic modeling via the joint discovery of latent topics and the different objective and subjective power hidden in every word; and 3) osLDA has lower computational complexity than supervised LDA, especially under an increasing number of topics.
Hanqi Wang, Fei Wu 0001, Weiming Lu 0001, Yi Yang 0001, Xi Li 0001, Xuelong Li 0001, Yueting Zhuang
IEEE Trans. Neural Networks Learn. Syst.3
2017 Cross-Lingual Entity Matching for Heterogeneous Online Wikis
Weiming Lu 0001, Baogang Wei
NLPCC1
2017 Boosting Collective Entity Linking via Type-Guided Semantic Embedding
Weiming Lu 0001, Haijiao Lu, Pengkun Ma, Zhenyu Zhang 0008, Baogang Wei
NLPCC1
2017 KeyphraseDS: Automatic generation of survey by exploiting keyphrase information
Shansong Yang, Weiming Lu 0001, Dezhi Yang, Xi Li 0001, Baogang Wei
Neurocomputing2
2017 Hybrid storage architecture and efficient MapReduce processing for unstructured data
abstract
As we are now entering the era of data deluge, how to efficiently manage these massive data is becoming a great challenge, especially for the exponentially growing unstructured data, which is far more than structured and semi-structured data. However, unstructured data is more complex for its variety. That is to say, different types of unstructured data have different file size, type and usage, which need different storage and processing for high efficiency. In this paper, we propose a hybrid storage architecture to store the pervasive unstructured data. This hybrid architecture integrates various kinds of data stores within a unified framework, where each type of unstructured data can find its suitable placement policy and it is transparent to users. In addition, we present several partitioning strategies based on the unified framework, which are beneficial to the MapReduce-based batch processing for these unstructured data. The experiments demonstrate that it is possible to build an efficient and smart system through the hybrid architecture and the partitioning strategies.
Weiming Lu 0001, Yaoguang Wang, Jingyuan Jiang, Yapeng Shen, Baogang Wei
Parallel Comput.1
2017 Regularized Deep Belief Network for Image Attribute Detection
abstract
In general, an image attribute is a human-nameable visual property that has a semantic connotation. Appropriate modeling of the intrinsic contextual correlations among attributes plays a fundamental role in attribute detection. In this paper, we consider image attribute detection from the perspective of regularized deep learning. In particular, we propose a regularized deep belief network (rDBN) to perform the image attribute detection task, which is composed of two parts: 1) a detection DBN (dDBN) that models the joint distribution of images and their corresponding attributes, which acts as an attribute detector and 2) a contextual restricted Boltzmann machine that explicitly models the correlations among attributes acting as a regularizer that restraints the output detection result given by the dDBN to meet the contextual prior of attributes. Furthermore, we propose an efficient fine-tuning scheme that can further optimize the performance of the dDBN by backpropagation. Experimental results show that the proposed rDBN obtains improvements over the state-of-the-art methods for attribute detection on the benchmark data sets.
Fei Wu 0001, Zhuhao Wang, Weiming Lu 0001, Xi Li 0001, Yi Yang 0001, Jiebo Luo 0001, Yueting Zhuang
IEEE Trans. Circuits Syst. Video Technol.3
2017 Bag-of-Discriminative-Words (BoDW) Representation via Topic Modeling
abstract
Many of the words in a given document either deliver facts (objective) or express opinions (subjective), respectively, depending on the topics they are involved in. For example, given a bunch of documents, the word “bug” assigned to the topic “order Hemiptera” apparently remarks one object (i.e., one kind of insects), while the same word assigned to the topic “software” probably conveys a negative opinion. Motivated by the intuitive assumption that different words have varying degrees of discriminative power in delivering the objective sense or the subjective sense with respect to their assigned topics, a model named as discriminatively objective-subjective LDA (dosLDA) is proposed in this paper. The essential idea underlying the proposed dosLDA is that a pair of objective and subjective selection variables are explicitly employed to encode the interplay between topics and discriminative power for the words in documents in a supervised manner. As a result, each document is appropriately represented as “bag-of-discriminativewords” (BoDW). The experiments reported on documents and images demonstrate that dosLDA not only performs competitively over traditional approaches in terms of topic modeling and document classification, but also has the ability to discern the discriminative power of each word in terms of its objective or subjective sense with respect to its assigned topic.
Yueting Zhuang, Hanqi Wang, Jun Xiao 0001, Fei Wu 0001, Yi Yang 0001, Weiming Lu 0001, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.6
2016 Diverse Image Captioning via GroupTalk
Zhuhao Wang, Fei Wu 0001, Weiming Lu 0001, Jun Xiao 0001, Xi Li 0001, Yueting Zhuang
IJCAI3
2016 D-Ocean: an unstructured data management system for data ocean environment
Yueting Zhuang, Yaoguang Wang, Jian Shao 0001, Ling Chen 0001, Weiming Lu 0001, Jianling Sun, Baogang Wei, Jiangqin Wu
Frontiers Comput. Sci.5
2016 Kernelized sparse hashing for scalable image retrieval
Yin Zhang 0006, Weiming Lu 0001, Yang Liu 0098, Fei Wu 0001
Neurocomputing2
2016 Amplifying scientific paper's abstract by leveraging data-weighted reconstruction
Shansong Yang, Weiming Lu 0001, Zhanjiang Zhang, Baogang Wei, Wenjia An
Inf. Process. Manag.2
2016 Aspect Learning for Multimedia Summarization via Nonparametric Bayesian
abstract
Summarization is desirable for efficient comprehension of an increasingly vast amount of data. A summary of multiple documents is a concise description of the main topic. Generally speaking, a topic delivers various aspects. For example, the natural disaster topic is likely to imply the aspects of casualties and rescue. Therefore, a good summary is expected to cover all the informative aspects of a topic in order to enhance both diversity and coverage of the topic. However, for the real-world data, the profile of aspects in a given topic (e.g., the number of the aspects as well as their appropriate describing sentences or images) is hardly specified in advance. To address this problem, this paper proposes an approach to learn the hidden aspects in the topics via a nonparametric Bayesian model for multimedia summarization, namely, aspect learning for multimedia summarization via nonparametric Bayesian (ALSNB). More specifically, we introduce the priors of beta-Bernoulli process and Dirichlet process into the traditional dictionary learning. As a result, the proposed approach is able to adaptively identify the particular aspects of an individual topic. The experimental results on several datasets for text summarization and image summarization show the superiority of the proposed ALSNB over other methods.
Fei Wu 0001, Hanyin Fang, Xi Li 0001, Siliang Tang, Weiming Lu 0001, Yi Yang 0001, Wenwu Zhu 0001, Yueting Zhuang
IEEE Trans. Circuits Syst. Video Technol.5
2015 Sketch the Storyline with CHARCOAL: A Non-Parametric Approach
Siliang Tang, Fei Wu 0001, Weiming Lu 0001, Zhongfei Zhang, Yueting Zhuang
IJCAI4
2015 Deep Compositional Cross-modal Learning to Rank via Local-Global Alignment
abstract
Cross-modal retrieval is a very hot research topic that is imperative to many applications involving multi-modal data. Discovering an appropriate representation for multi-modal data and learning a ranking function are essential to boost the cross-media retrieval. Motivated by the assumption that a compositional cross-modal semantic representation (pairs of images and text) is more attractive for cross-modal ranking, this paper exploits the existing image-text databases to optimize a ranking function for cross-modal retrieval, called deep compositional cross-modal learning to rank (C2MLR). In this paper, C2MLR considers learning a multi-modal embedding from the perspective of optimizing a pairwise ranking problem while enhancing both local alignment and global alignment. In particular, the local alignment (i.e., the alignment of visual objects and textual words) and the global alignment (i.e., the image-level and sentence-level alignment) are collaboratively utilized to learn the multi-modal embedding common space in a max-margin learning to rank manner. The experiments demonstrate the superiority of our proposed C2MLR due to its nature of multi-modal compositional embedding.
Xinyang Jiang, Fei Wu 0001, Xi Li 0001, Zhou Zhao 0001, Weiming Lu 0001, Siliang Tang, Yueting Zhuang
ACM Multimedia5
2015 Short Text Understanding by Leveraging Knowledge into Topic Model
abstract
Shansong Yang, Weiming Lu, Dezhi Yang, Liang Yao, Baogang Wei. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Shansong Yang, Weiming Lu 0001, Dezhi Yang, Baogang Wei
HLT-NAACL2
2015 Taxonomy Induction from Chinese Encyclopedias by Combinatorial Optimization
Weiming Lu 0001, Renjie Lou, Zhenyu Zhang 0008, Shansong Yang, Baogang Wei
NLPCC1
2015 Mining RDF from Tables in Chinese Encyclopedias
abstract
Web tables understanding has recently attracted a number of studies. However, many works focus on the tables in English, because they usually need the help of knowledge bases, while the existing knowledge bases such as DBpedia, YAGO, Freebase and Probase mainly contain knowledge in English. In this paper, we focus on the RDF triples extraction from tables in Chinese encyclopedias. Firstly, we constructed a Chinese knowledge base through taxonomy mining and class attribute mining. Then, with the help of our knowledge base, we extracted triples from tables through column scoring, table classification and RDF extraction. In our experiments, we practically implemented our approach in 6,618,544 articles from Hudong Baike with 764,292 tables, and extracted about 1,053,407 unique and new RDF triples with an estimated accuracy of \(90.2\%\) , which outperforms other similar works.
Weiming Lu 0001, Zhenyu Zhang 0008, Renjie Lou, Shansong Yang, Baogang Wei
NLPCC1
2015 Improving MapReduce Performance with Partial Speculative Execution
Yaoguang Wang, Weiming Lu 0001, Renjie Lou, Baogang Wei
J. Grid Comput.2
2015 Topic aspect-oriented summarization via group selection
Hanyin Fang, Weiming Lu 0001, Fei Wu 0001, Yin Zhang 0006, Xindi Shang, Jian Shao 0001, Yueting Zhuang
Neurocomputing2
2015 The classification of multi-modal data with hidden conditional random field
Xinyang Jiang, Fei Wu 0001, Yin Zhang 0006, Siliang Tang, Weiming Lu 0001, Yueting Zhuang
Pattern Recognit. Lett.5
2015 Cross-Modal Learning to Rank via Latent Joint Representation
abstract
Cross-modal ranking is a research topic that is imperative to many applications involving multimodal data. Discovering a joint representation for multimodal data and learning a ranking function are essential in order to boost the cross-media retrieval (i.e., image-query-text or text-query-image). In this paper, we propose an approach to discover the latent joint representation of pairs of multimodal data (e.g., pairs of an image query and a text document) via a conditional random field and structural learning in a listwise ranking manner. We call this approach cross-modal learning to rank via latent joint representation (CML²R). In CML²R, the correlations between multimodal data are captured in terms of their sharing hidden variables (e.g., topics), and a hidden-topic-driven discriminative ranking function is learned in a listwise ranking manner. The experiments show that the proposed approach achieves a good performance in cross-media retrieval and meanwhile has the capability to learn the discriminative representation of multimodal data.
Fei Wu 0001, Xinyang Jiang, Xi Li 0001, Siliang Tang, Weiming Lu 0001, Zhongfei Zhang, Yueting Zhuang
IEEE Trans. Image Process.5
2015 Structured Visual Feature Learning for Classification via Supervised Probabilistic Tensor Factorization
abstract
In this paper, structured visual feature learning aims at exploiting the intrinsic structural properties of mutually correlated multimedia collections (e.g., video frames or facial images) to learn a more effective feature representation for multimedia data classification. We pose structured visual feature learning as a problem of supervised tensor factorization (STF), which is capable of effectively learning multi-view visual features from structural tensorial multimedia data. In mathematics , STF is formulated as a joint optimization framework of probabilistic inference and$\epsilon $-insensitive support vector regression. As a result, the feature representation obtained by STF not only preserves the intrinsic multi-view structural information on tensorial multimedia data, but also includes the discriminative information derived from the max-margin learning process. Using the learned discriminative visual features, we conduct a set of multimedia classification experiments on several challenging datasets, including images and videos, which demonstrate the effectiveness of our method.
Xu Tan 0003, Fei Wu 0001, Xi Li 0001, Siliang Tang, Weiming Lu 0001, Yueting Zhuang
IEEE Trans. Multim.5
2014 Geo-informative discriminative image representation by semi-supervised hierarchical topic modeling
abstract
Nowadays, the prevalence of sharing tourist photos to online communities has created an increasing demand for mining discriminative architecture aspects from historic landmarks. Some previous researches have demonstrated that topic models could discover discriminative features represented by meaningful visual-topics. However, they seldom exploited the indicative function of geo-tags and the hierarchy in architecture characteristics. In order to utilize this information, we proposed a semi-supervised hierarchical topic modeling approach (namely, shTM). In our approach, every image could be represented by a probability distribution over selected geo-related visual-topics from a partly randomized topic tree. We evaluated our approach on a real-world dataset with over 26 thousand geo-informative photos from Flickr. Experiments show that shTM topics could reveal more discriminative aspects of a specific architecture than other well-known image features, such as HOG and SIFT, on the tasks of automatic photo categorization and geographical information retrieval.
Zijian Li 0002, Siliang Tang, Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang
ICME4
2014 Learning Multimodal Neural Network with Ranking Examples
abstract
To support cross-modal information retrieval, cross-modal learning to rank approaches utilize ranking examples (e.g., an example may be a text query and its corresponding ranked images) to learn appropriate ranking (similarity) function. However, the fact that each modality is represented with intrinsically different low-level features hinders these approaches from better reducing the heterogeneity-gap between the modalities and thus giving satisfactory retrieval results. In this paper, we consider learning with neural networks, from the perspective of optimizing the listwise ranking loss of the cross-modal ranking examples. The proposed model, named Cross-Modal Ranking Neural Network (CMRNN), benefits from the advance of both neural networks on learning high-level semantics and learning to rank techniques on learning ranking function, such that the learned cross-modal ranking function is implicitly embedded in the learned high-level representation for data objects with different modalities (e.g., text and imagery) to perform cross-modal retrieval directly. We compare CMRNN to existing state-of-the-art cross-modal ranking methods on two datasets and show that it achieves a better performance.
Fei Wu 0001, Xi Li 0001, Yin Zhang 0006, Weiming Lu 0001, Yueting Zhuang
ACM Multimedia5
2014 Multiple kernel learning with NOn-conVex group spArsity
Weiming Lu 0001, Fei Wu 0001, Yueting Zhuang
J. Vis. Commun. Image Represent.2
2013 Digital Library Engine: Adapting Digital Library for Cloud Computing
abstract
With the rapid growth of digital libraries, more data and smart services are involved. People come to recognize the importance of digital libraries and the convenience they might bring to the society. However, the cost of owning a digital library is quite high, and many institutions do not have the ability to run and maintain a digital library by themselves, especially for massive data and complex services which require lots of storage and computing resources. In this paper, we proposed the Digital Library Engine, which aims to provide a new Platform as a Service for fast developing and deploying digital libraries in cloud. To the best of our knowledge, it is the first work to create a PaaS system for digital libraries. With the help of Digital Library Engine, institutions only need to develop some service bundles, which can be deployed in the engine, and then their own digital libraries could be running well with features of scalability, reliability, security, extensibility, availability and manageability. The practice in CADAL and the experiments demonstrate the feasibility and efficiency of our engine.
Weiming Lu 0001, Liangju Zheng, Jian Shao 0001, Baogang Wei, Yueting Zhuang
IEEE CLOUD1
2013 Supervised Coupled Dictionary Learning with Group Structures for Multi-modal Retrieval
abstract
A better similarity mapping function across heterogeneous high-dimensional features is very desirable for many applications involving multi-modal data. In this paper, we introduce coupled dictionary learning (DL) into supervised sparse coding for multi-modal (cross-media) retrieval. We call this Supervised coupled dictionary learning with group structures for Multi-Modal retrieval (SliM2). SliM2 formulates the multi-modal mapping as a constrained dictionary learning problem. By utilizing the intrinsic power of DL to deal with the heterogeneous features, SliM2 extends unimodal DL to multi-modal DL. Moreover, the label information is employed in SliM2 to discover the shared structure inside intra-modality within the same class by a mixed norm (i.e., `l1/l2`-norm). As a result, the multimodal retrieval is conducted via a set of jointly learned mapping functions across multi-modal data. The experimental results show the effectiveness of our proposed model when applied to cross-media retrieval.
Yueting Zhuang, Fei Wu 0001, Yin Zhang 0006, Weiming Lu 0001
AAAI5
2012 Transactional Multi-row Access Guarantee in the Key-Value Store
abstract
The emergence of Cloud Computing and Big Data drives the development of novel data stores named NoSQL. A mass of data stores are developed and the most are key-value stores, where the stores are partitioned with keys and a key can identify a row uniquely. However, the requirement for efficiency and scalability makes them only provide the single-row atomic access. But in the Big Data era, more and more applications built on the key-value stores need transactional functionality across multiple rows. So, it is natural to implement a multi-row transaction management for key-value stores. In this paper, we implement a transaction processing system (TrasPS) which guarantees the transactional multi-row access from the application client to the key-value store in our unstructured data management system (UDMS). We also provide fault tolerance and recovery for the transactions. The implementation and experiments in our UDMS show that TrasPS can provide scalable multi-row access functionality at a very low overhead.
Yaoguang Wang, Weiming Lu 0001, Baogang Wei
CLUSTER2
2012 HAaaS: Towards Highly Available Distributed Systems
abstract
High availability is a valuable property in distributed systems. The master-slave model is used wildly in data management systems for high performance. However, many master-slave systems still have SPOF (Single Point of Failure) for the single master node. We exploit a generalized solution to meet several common use cases for different master-slave systems. The solution makes the high availability as a service (HAaaS), which uses a shared storage infrastructure to make the master stateless and provides an automatic fail over of high-availability service. We deploy the HAaaS in many master-slave subsystems in our unstructured data management system (UDMS) to make the UDMS highly available. The experiments demonstrate the feasibility and efficiency of our solution.
Yaoguang Wang, Weiming Lu 0001, Baogang Wei
CLUSTER2
2012 A unified framework for web video topic discovery and visualization
Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang
Pattern Recognit. Lett.3
2011 Efficient shape matching for Chinese calligraphic character retrieval
abstract
An efficient search method is desired for calligraphic characters due to the explosive growth of calligraphy works in digital libraries. However, traditional optical character recognition (OCR) and handwritten character recognition (HCR) technologies are not suitable for calligraphic character retrieval. In this paper, a novel shape descriptor called SC-HoG is proposed by integrating global and local features for more discriminability, where a gradient descent algorithm is used to learn the optimal combining parameter. Then two efficient methods, keypoint-based method and locality sensitive hashing (LSH) based method, are proposed to accelerate the retrieval by reducing the feature set and converting the feature set to a feature vector. Finally, a re-ranking method is described for practicability. The approach filters query-dissimilar characters using the LSH-based method to obtain candidates first, and then re-ranks the candidates using the keypoint- or sample-based method. Experimental results demonstrate that our approaches are effective and efficient for calligraphic character retrieval.
Weiming Lu 0001, Jiangqin Wu, Baogang Wei, Yueting Zhuang
J. Zhejiang Univ. Sci. C1
2010 CMSOF: a structured data organization framework for scanned Chinese medicine books in digital libraries
abstract
Organizing unstructured information from books into a well-defined structure is a significant challenge in digital libraries. Most digital libraries can provide only search services at the granularity of books and few libraries allow books to be accessed at the granularity of chapters, as manually constructing directory information for books is time-consuming. Extracting structured data from scanned books thus remains an urgent and important work. In this paper, we propose a novel structured data organization framework called CMSOF to organize scanned data automatically, and apply it to a Chinese medicine digital library. In the framework, image blocks and text blocks on the scanned page of books are separated based on the gray histogram projection method or a hybrid method of region growth and the Ada-Boosting classifier at first, and then the text structure is obtained from text blocks by text size and font type recognition. Finally, image blocks and structured OCRed text are correlated at the semantic level. By integrating the structured data into a Chinese medicine information system (CMIS), we can organize the Chinese medicine books well and users can access the books with flexibility, which indicates that CMSOF is an efficient framework to organize books mixed with images and text.
Baogang Wei, Weiming Lu 0001, Yueting Zhuang
J. Zhejiang Univ. Sci. C4
2009 Latent Style Model: Discovering writing styles for calligraphy works
Yueting Zhuang, Weiming Lu 0001, Jiangqin Wu
J. Vis. Commun. Image Represent.2
2009 Discovering calligraphy style relationships by Supervised Learning Weighted Random Walk Model
Weiming Lu 0001, Yueting Zhuang, Jiangqin Wu
Multim. Syst.1