EDBT 2026 Demo / reviewers in the wild / expert
Can Ma
dblp:62/3794
· DBLP profile ↗
50ranked-venue papers
3as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 18 since 2021Systems, architecture and hardware · 7 · 1 first-author · 1 since 2021Computer networks · 4 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Communication-efficient personalized federal graph learning via low-rank decomposition
Ruyue Liu, Rong Yin 0001, Xiangzhen Bo, Xiaoshuai Hao, Xingrui Zhou, Yong Liu 0018, Jinwen Zhong, Can Ma, Weiping Wang 0005 |
Pattern Recognit. | 9 |
| 2026 | EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration
Daiqing Wu, Dongbao Yang, Can Ma, Yu Zhou 0015 |
Pattern Recognit. | 3 |
| 2026 | Resolving sentiment discrepancy for multimodal sentiment detection via semantics completion and decomposition
Daiqing Wu, Dongbao Yang, Huawen Shen, Can Ma, Yu Zhou 0015 |
Pattern Recognit. | 4 |
| 2025 | Union Is Strength! Unite the Power of LLMs and MLLMs for Chart Question AnsweringabstractChart Question Answering (CQA) requires models to perform chart perception and reasoning. Recent studies driven by Large Language Models (LLMs) have dominated CQA. These include employing more cognitively capable LLMs for indirectly reasoning over transformed charts, i.e., tables, and directly perceiving charts utilizing Multimodal Large Language Models (MLLMs) with a wider perceptual range. Yet, they often encounter bottlenecks due to the limitation of the receptive field of LLMs and the fragility of the complex reasoning of some MLLMs. To unite the strengths of LLMs and MLLMs to complement each other's limitations, we propose Synergy, a framework that unites the power of both models for CQA. Synergy first unites the chart with a table as the augmented perceptual signal. Next, it unites LLMs and MLLMs, scheduling the former to decompose a question into subquestions and the latter to answer these by perceiving the chart. Lastly, it operates LLMs to summarize the subquestion-answer pairs to refine the final answer. Extensive experimental results on popular CharQA and PlotQA benchmarks reveal that, with the power of union, Synergy outperforms strong competitors and achieves superior boosts over naive MLLMs by uniting them with a smaller LLM. Jiapeng Liu 0006, Shihao Rao, Xiyan Gao, Weixin Guan, Can Ma |
AAAI | 7 |
| 2025 | Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal CluesabstractVideo text-based visual question answering (TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. Inspired by the development of TextVQA in image domain, existing Video TextVQA approaches leverage a language model (e.g. T5) to process text-rich multiple frames and generate answers auto-regressively. Nevertheless, the spatio-temporal relationships among visual entities (including scene text and objects) will be disrupted and models are susceptible to interference from unrelated information, resulting in irrational reasoning and inaccurate answering. To tackle these challenges, we propose the TEA (stands for "Track the Answer'') method that better extends the generative TextVQA framework from image to video. TEA recovers the spatio-temporal relationships in a complementary way and incorporates OCR-aware clues to enhance the quality of reasoning questions. Extensive experiments on several public Video TextVQA datasets validate the effectiveness and generalization of our framework. TEA outperforms existing TextVQA methods, video-language pretraining methods and video large language models by great margins. The code will be publicly released. Gangyan Zeng, Huawen Shen, Daiqing Wu, Yu Zhou 0015, Can Ma |
AAAI | 6 |
| 2025 | Linguistics-aware Masked Image Modeling for Self-supervised Scene Text RecognitionabstractText images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates contextual and semantic elements. In scenarios with degraded visual quality, linguistic patterns serve as crucial supplements for comprehension, highlighting the necessity of integrating both aspects for robust scene text recognition (STR). Contemporary STR approaches often use language models or semantic reasoning modules to capture linguistic features, typically requiring large-scale annotated datasets. Self-supervised learning, which lacks annotations, presents challenges in disentangling linguistic features related to the global context. Typically, sequence contrastive learning emphasizes the alignment of local features, while masked image modeling (MIM) tends to exploit local structures to reconstruct visual patterns, resulting in limited linguistic knowledge. In this paper, we propose a Linguistics-aware Masked Image Modeling (LMIM) approach, which channels the linguistic information into the decoding process of MIM through a separate branch. Specifically, we design a linguistics alignment module to extract vision-independent features as linguistic guidance using inputs with different visual appearances. As features extend beyond mere visual structures, LMIM must consider the global context to achieve reconstruction. Extensive experiments on various benchmarks quantitatively demonstrate our state-of-the-art performance, and attention visualizations qualitatively show the simultaneous capture of both visual and linguistic information. The code is available at https: //github.com/zhangyifei01/LMIM. Yifei Zhang 0005, Yu Zhou 0015, Can Ma, Xiangyang Ji |
CVPR | 6 |
| 2025 | PACM: Position-Aware Cross-Modality Decoder for Handwritten Mathematical Expression Recognition
Zhijie Shen, Can Ma, Yaqiang Wu, Yu Zhou 0015 |
ICDAR (1) | 4 |
| 2025 | An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception CapabilityabstractThe advancements in Multimodal Large Language Models (MLLMs) have enabled various multimodal tasks to be addressed under a zero-shot paradigm. This paradigm sidesteps the cost of model fine-tuning, emerging as a dominant trend in practical application. Nevertheless, Multimodal Sentiment Analysis (MSA), a pivotal challenge in the quest for general artificial intelligence, fails to accommodate this convenience. The zero-shot paradigm exhibits undesirable performance on MSA, casting doubt on whether MLLMs can perceive sentiments as competent as supervised models. By extending the zero-shot paradigm to In-Context Learning (ICL) and conducting an in-depth study on configuring demonstrations, we validate that MLLMs indeed possess such capability. Specifically, three key factors that cover demonstrations' retrieval, presentation, and distribution are comprehensively investigated and optimized. A sentimental predictive bias inherent in MLLMs is also discovered and later effectively counteracted. By complementing each other, the devised strategies for three factors result in average accuracy improvements of 15.9% on six MSA datasets against the zero-shot paradigm and 11.2% against the random ICL baseline. Daiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma, Yu Zhou 0015 |
ICML | 4 |
| 2025 | Improving Mathematical Reasoning Abilities of Small Language Models via Key-Point-Driven DistillationabstractLarge Language Models (LLMs) excel in mathematical reasoning due to their extensive parameters and training data, but their high computational demands hinder deployment. Distilling LLM reasoning into Smaller Language Models (SLMs, ≤ 1B parameters) is a potential solution, yet these models often struggle with calculation and semantic errors. Previous work introduced Program-of-Thought Distillation (PoTD) to reduce calculation mistakes. To tackle semantic errors, we propose Key-Point-Driven Mathematical Reasoning Distillation (KPDD), which improves SLM reasoning by splitting the problem-solving process into Key Points Extraction and Step-by-Step Solution. KPDD includes KPDD-CoT, generating Chain-of-Thought rationales, and KPDD-PoT, producing Program-of-Thought rationales. Experiments show KPDD-CoT enhances reasoning capabilities, while KPDD-PoT achieves state-of-the-art performance in mathematical tasks, effectively reducing misunderstanding errors and promoting the deployment of efficient, capable SLMs. Xunyu Zhu, Jian Li 0040, Rong Yin 0001, Can Ma, Weiping Wang 0005 |
IJCNN | 4 |
| 2025 | Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven DistillationabstractLarge Language Models (LLMs) demonstrate exceptional reasoning capabilities, often achieving state-of-the-art performance in various tasks. However, their substantial computational and memory demands, due to billions of parameters, hinder deployment in resource-constrained environments. A promising solution is knowledge distillation, where LLMs transfer reasoning capabilities to Small Language Models (SLMs, ≤ 1B parameters), enabling wider deployment on low-resource devices. Existing methods primarily focus on generating high-quality reasoning rationales for distillation datasets but often neglect the critical role of data quantity and quality. To address these challenges, we propose a Feedback-Driven Distillation (FDD) framework to enhance SLMs’ mathematical reasoning capabilities. In the initialization stage, a distillation dataset is constructed by prompting LLMs to pair mathematical problems with corresponding reasoning rationales. We classify problems into easy and hard categories based on SLM performance. For easy problems, LLMs generate more complex variations, while for hard problems, new questions of similar complexity are synthesized. In addition, we propose a multi-round distillation paradigm to iteratively enrich the distillation datasets, thereby progressively improving the mathematical reasoning abilities of SLMs. Experimental results demonstrate that our method can make SLMs achieve SOTA mathematical reasoning performance. Xunyu Zhu, Jian Li 0040, Rong Yin 0001, Can Ma, Weiping Wang 0005 |
IJCNN | 4 |
| 2025 | Gather and Trace: Rethinking Video TextVQA from an Instance-oriented PerspectiveabstractVideo text-based visual question answering (Video TextVQA) aims to answer questions by explicitly reading and reasoning about the text involved in a video. Most works in this field follow a frame-level framework which suffers from redundant text entities and implicit relation modeling, resulting in limitations in both accuracy and efficiency. In this paper, we rethink the Video TextVQA task from an instance-oriented perspective and propose a novel model termed GAT (Gather and Trace). First, to obtain accurate reading result for each video text instance, a context-aggregated instance gathering module is designed to integrate the visual appearance, layout characteristics, and textual contents of the related entities into a unified textual representation. Then, to capture dynamic evolution of text in the video flow, an instance-focused trajectory tracing module is utilized to establish spatio-temporal relationships between instances and infer the final answer. Extensive experiments on several public Video TextVQA datasets validate the effectiveness and generalization of our framework. GAT outperforms existing Video TextVQA methods, video-language pretraining methods, and video large language models in both accuracy and inference speed. Notably, GAT surpasses the previous state-of-the-art Video TextVQA methods by 3.86% in accuracy and achieves ten times of faster inference speed than video large language models. The source code is available at https://github.com/zhangyan-ucas/GAT. Gangyan Zeng, Daiqing Wu, Huawen Shen, Binbin Li 0003, Yu Zhou 0015, Can Ma, Xiaojun Bi 0002 |
ACM Multimedia | 7 |
| 2025 | SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed GraphsabstractLarge-scale pre-trained models have revolutionized Natural Language Processing (NLP) and Computer Vision (CV), showcasing remarkable cross-domain generalization abilities. However, in graph learning, models are typically trained on individual graph datasets, limiting their capacity to transfer knowledge across different graphs and tasks. This approach also heavily relies on large volumes of annotated data, which presents a significant challenge in resource-constrained settings. Unlike NLP and CV, graph-structured data presents unique challenges due to its inherent heterogeneity, including domain-specific feature spaces and structural diversity across various applications. To address these challenges, we propose a novel structure-aware self-supervised learning method for Text-Attributed Graphs (SSTAG). By leveraging text as a unified representation medium for graph learning, SSTAG bridges the gap between the semantic reasoning of Large Language Models (LLMs) and the structural modeling capabilities of Graph Neural Networks (GNNs). Our approach introduces a dual knowledge distillation framework that co-distills both LLMs and GNNs into structure-aware multilayer perceptrons (MLPs), enhancing the scalability of large-scale TAGs. Additionally, we introduce an in-memory mechanism that stores typical graph representations, aligning them with memory anchors in an in-memory repository to integrate invariant knowledge, thereby improving the model’s generalization ability. Extensive experiments demonstrate that SSTAG outperforms state-of-the-art models on cross-domain transfer learning tasks, achieves exceptional scalability, and reduces inference costs while maintaining competitive performance. Ruyue Liu, Rong Yin 0001, Xiangzhen Bo, Xiaoshuai Hao, Yong Liu 0018, Jinwen Zhong, Can Ma, Weiping Wang 0005 |
NeurIPS | 7 |
| 2025 | Enhancing Cybersecurity in the Big Data Era: A GA-Optimized Fuzzy Clustering ApproachabstractResearch Highlights • This study proposes a novel approach for enhancing cybersecurity through the integration of GA-AFCM. This method significantly improves the accuracy, efficiency, and adaptability of IDS. • The GA-AFCM technique demonstrates superior performance compared to conventional methods such as K-Means, MKKM-IC, Density Peaks, and GMM. • The proposed method effectively addresses the challenges of security of information in the period of Big Data, achieving the highest detection rate while significantly reducing false positives, thereby enhancing overall system reliability and efficiency. Cybersecurity includes protecting computer networks and systems from unauthorized access, harm, and fraud, employing various techniques and technologies such as barriers, antivirus software, and cryptography. Regular system updates, employee training, and adherence to best practices are crucial for maintaining confidentiality and ensuring reliable IT services in both corporate and public sectors. This paper introduces a GA-AFCM technique, which enhances intrusion detection and cybersecurity tasks by combining the strengths of Genetic Algorithms and Adaptive Fuzzy C-Means Clustering. The study began with data collection and preprocessing using Z-score normalization, followed by feature extraction through Linear Discrimination Analysis (LDA). The GA-AFCM approach was compared with traditional methods such as K-Means, Density Peaks, GMM, and MKKM-IC. The results demonstrate the TPR (91%), FPR (4%) precision (83.56%), accuracy (95.6%), and F1-score (87%) are used to examining and interpreting quickly and dynamically generated data streams efficiently solved by the proposed approach. The GA-AFCM method significantly enhances detection rates to over 95% while substantially reducing false positives, establishing it as a robust solution for cybersecurity in the big data era. Tieguang Xu, Can Ma, Zhaozhao Su, Jingqiong Su, Zhiming Ma, Jianzhen Wang, Zhaolong Yang |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |
| 2025 | Multi-Modal Molecular Representation Learning via Structure AwarenessabstractAccurate extraction of molecular representations is a critical step in the drug discovery process. In recent years, significant progress has been made in molecular representation learning methods, among which multi-modal molecular representation methods based on images, and 2D/3D topologies have become increasingly mainstream. However, existing these multi-modal approaches often directly fuse information from different modalities, overlooking the potential of intermodal interactions and failing to adequately capture the complex higher-order relationships and invariant features between molecules. To overcome these challenges, we propose a structure-awareness-based multi-modal self-supervised molecular representation pre-training framework (MMSA) designed to enhance molecular graph representations by leveraging invariant knowledge between molecules. The framework consists of two main modules: the multi-modal molecular representation learning module and the structure-awareness module. The multi-modal molecular representation learning module collaboratively processes information from different modalities of the same molecule to overcome intermodal differences and generate a unified molecular embedding. Subsequently, the structure-awareness module enhances the molecular representation by constructing a hypergraph structure to model higher-order correlations between molecules. This module also introduces a memory mechanism for storing typical molecular representations, aligning them with memory anchors in the memory bank to integrate invariant knowledge, thereby improving the model's generalization ability. Compared to existing multi-modal approaches, MMSA can be seamlessly integrated with any graph-based method and supports multiple molecular data modalities, ensuring both versatility and compatibility. Extensive experiments have demonstrated the effectiveness of MMSA, which achieves state-of-the-art performance on the MoleculeNet benchmark, with average ROC-AUC improvements ranging from 1.8% to 9.6% over baseline methods. Rong Yin 0001, Ruyue Liu, Xiaoshuai Hao, Xingrui Zhou, Yong Liu 0018, Can Ma, Weiping Wang 0005 |
IEEE Trans. Image Process. | 6 |
| 2025 | AS-GCL: Asymmetric Spectral Augmentation on Graph Contrastive LearningabstractGraph Contrastive Learning (GCL) has emerged as the foremost approach for self-supervised learning on graph-structured data. GCL reduces reliance on labeled data by learning robust representations from various augmented views. However, existing GCL methods typically depend on consistent stochastic augmentations, which overlook their impact on the intrinsic structure of the spectral domain, thereby limiting the model's ability to generalize effectively. To address these limitations, we propose a novel paradigm called AS-GCL that incorporates asymmetric spectral augmentation for graph contrastive learning. A typical GCL framework consists of three key components: graph data augmentation, view encoding, and contrastive loss. Our method introduces significant enhancements to each of these components. Specifically, for data augmentation, we apply spectral-based augmentation to minimize spectral variations, strengthen structural invariance, and reduce noise. With respect to encoding, we employ parameter-sharing encoders with distinct diffusion operators to generate diverse, noise-resistant graph views. For contrastive loss, we introduce an upper-bound loss function that promotes generalization by maintaining a balanced distribution of intra- and inter-class distance. To our knowledge, we are the first to encode augmentation views of the spectral domain using asymmetric encoders. Extensive experiments on eight benchmark datasets across various node-level tasks demonstrate the advantages of the proposed method. Ruyue Liu, Rong Yin 0001, Yong Liu 0018, Xiaoshuai Hao, Haichao Shi, Can Ma, Weiping Wang 0005 |
IEEE Trans. Multim. | 6 |
| 2025 | TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language ModelabstractExisting scene text spotters are designed to locate and transcribe texts from images. However, it is challenging for a spotter to achieve precise detection and recognition of scene texts simultaneously. Inspired by the glimpse-focus spotting pipeline of human beings and impressive performances of Pre-trained Language Models (PLMs) on visual tasks, we ask: (1) “Can machines spot texts without precise detection just like human beings?”, and if yes, (2) “Is text block another alternative for scene text spotting other than word or character?” To this end, our proposed scene text spotter leverages advanced PLMs to enhance performance without fine-grained detection. Specifically, we first use a simple detector for block-level text detection to obtain rough positional information. Then, we fine-tune a PLM using a large-scale OCR dataset to achieve accurate recognition. Benefiting from the comprehensive language knowledge gained during the pre-training phase, the PLM-based recognition module effectively handles complex scenarios, including multi-line, reversed, occluded, and incomplete-detection texts. Taking advantage of the fine-tuned language model on scene recognition benchmarks and the paradigm of text block detection, extensive experiments demonstrate the superior performance of our scene text spotter across multiple public benchmarks. Additionally, we attempt to spot texts directly from an entire scene image to demonstrate the potential of PLMs, even Large Language Models (LLMs). Jiahao Lyu 0002, Gangyan Zeng, Enze Xie, Wei Wang 0315, Can Ma, Yu Zhou 0015 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | Free your mouse! Command Large Language Models to Generate Code to Format Word DocumentsabstractRecently, LLMs have significantly improved code generation, making it increasingly accessible to users.As a result, LLM-powered code generation applications have sprung up, vastly boosting user productivity.This paper mainly explores how to improve the efficiency and experience of users in formatting the document.Specifically, we propose an automatic document formatting method, TEXT-TO-FORMAT, which is driven by various prompting strategies.TEXT-TO-FORMAT takes the user's formatting instructions and then generates code that can be run in Microsoft Word to format the content in a document.Further, to evaluate automatic document formatting approaches and advance the document formatting task, we build an evaluation specification including a high-quality dataset DOCFORMEVAL, a code runtime environment, and evaluation metrics.Extensive experimental results on DOCFORMEVAL reveal that the prompting strategy's effect positively correlates with how much knowledge it introduces related to document formatting task.We believe the constructed DOCFORMEVAL and the exploration about TEXT-TO-FORMAT can help developers build more intelligent tools for automatic document formatting, especially in offline scenarios, where the data privacy is the top priority 1 . Shihao Rao, Jiapeng Liu 0006, Weixin Guan, Xiyan Gao, Bing Lim, Can Ma |
EMNLP | 7 |
| 2024 | Segment then Match: Find the Carrier before Reasoning in Scene-Text VQAabstractText-based Visual Question Answering (TextVQA) requires models to answer questions about the scene text in images by reasoning the context between the scene text and the question. Previous works demonstrated that clustering the scene text could help the model understand the context between different scene texts in the image. However, these methods cluster scene text solely based on spatial information, resulting in close scene text being grouped together even in the absence of semantic contextual relationships. In order to solve the above problem, we proposed a Segment then Match method. Specifically, we propose an OCR-carrier Segmentation and Matching module to segment texts and carriers in the scene image and help all OCR texts find the carrier that belongs to them. Then, we propose a Hierarchical Visual Feature Fusion module to facilitate semantic relevance judgment of OCR text from multiple visual perspectives, thereby aiding the answer reasoning process. Our proposed method outperforms state-of-the-art methods by 3.65% and 3.31% on TextVQA and ST-VQA datasets, respectively. Extensive experiments validate the effectiveness of our method. Chengyang Fang, Jiapeng Liu 0006, Dayong Hu, Can Ma |
ICASSP | 6 |
| 2024 | Prompting Large Language Models with Fine-Grained Visual Relations from Scene Graph for Visual Question AnsweringabstractVisual Question Answering (VQA) is a task that requires models to comprehend both questions and images. An increasing number of works are leveraging the strong reasoning capabilities of Large Language Models (LLMs) to address VQA. These methods typically utilize image captions as visual text description to aid LLMs in comprehending images. However, these captions often overlooking the relations of fine-grained objects, which will limit the reasoning capability of LLMs. In this paper, we present PFVR, a modular framework that Prompts LLMs with Fine-grained Visual Relationships for VQA. PFVR primarily consists of an answer-guided generation module (AGG) and a question-guided filtering module (QGF). The two modules can combine to extract the fine-grained visual relations from scene graph, which will finally serve as crucial context for LLMs to comprehend the image. Extensive experiments conducted on the popular VQA dataset, GQA, confirm PFVR achieves state-of-the-art results compared to other strong VQA competitors, demonstrating its exceptional effectiveness. Jiapeng Liu 0006, Chengyang Fang, Dayong Hu, Can Ma |
ICASSP | 6 |
| 2024 | Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text PairsabstractVisual emotion recognition (VER) is a longstanding field that has garnered increasing attention with the advancement of deep neural networks. Although recent studies have achieved notable improvements by leveraging the knowledge embedded within pre-trained visual models, the lack of direct association between factual-level features and emotional categories, called the ''affective gap'', limits the applicability of pre-training knowledge for VER tasks. On the contrary, the explicit emotional expression and high information density in textual modality eliminate the ''affective gap''. Therefore, we propose borrowing the knowledge from the pre-trained textual model to enhance the emotional perception of pre-trained visual models. We focus on the factual and emotional connections between images and texts in noisy social media data, and propose Partitioned Adaptive Contrastive Learning (PACL) to leverage these connections. Specifically, we manage to separate different types of samples and devise distinct contrastive learning strategies for each type. By dynamically constructing negative and positive pairs, we fully exploit the potential of noisy samples. Through comprehensive experiments, we demonstrate that bridging the "affective gap'' significantly improves the performance of various pre-trained visual models in downstream emotion-related tasks. Our code is released on https://github.com/wdqqdw/PACL. Daiqing Wu, Dongbao Yang, Yu Zhou 0015, Can Ma |
ACM Multimedia | 4 |
| 2024 | Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and FusionabstractAs posts on social media increase rapidly, analyzing the sentiments embedded in image-text pairs has become a popular research topic in recent years. Although existing works achieve impressive accomplishments in simultaneously harnessing image and text information, they lack the considerations of possible low-quality and missing modalities. In real-world applications, these issues might frequently occur, leading to urgent needs for models capable of predicting sentiment robustly. Therefore, we propose a Distribution-based feature Recovery and Fusion (DRF) method for robust multimodal sentiment analysis of image-text pairs. Specifically, we maintain a feature queue for each modality to approximate their feature distributions, through which we can simultaneously handle low-quality and missing modalities in a unified framework. For low-quality modalities, we reduce their contributions to the fusion by quantitatively estimating modality qualities based on the distributions. For missing modalities, we build inter-modal mapping relationships supervised by samples and distributions, thereby recovering the missing modalities from available ones. In experiments, two disruption strategies that corrupt and discard some modalities in samples are adopted to mimic the low-quality and missing modalities in various real-world scenarios. Through comprehensive experiments on three publicly available image-text datasets, we demonstrate the universal improvements of DRF compared to SOTA methods under both two strategies, validating its effectiveness in robust multimodal sentiment analysis. Daiqing Wu, Dongbao Yang, Yu Zhou 0015, Can Ma |
ACM Multimedia | 4 |
| 2024 | Show Exemplars and Tell Me What You See: In-Context Learning with Frozen Large Language Models for TextVQA
Gangyan Zeng, Huawen Shen, Can Ma, Yu Zhou 0015 |
PRCV (7) | 4 |
| 2024 | Distilling mathematical reasoning capabilities into Small Language Models
Xunyu Zhu, Jian Li 0040, Yong Liu 0018, Can Ma, Weiping Wang 0005 |
Neural Networks | 4 |
| 2024 | Unifying Structured Data as Graph for Data-to-Text Pre-TrainingabstractAbstract Data-to-text (D2T) generation aims to transform structured data into natural language text. Data-to-text pre-training has proved to be powerful in enhancing D2T generation and yields impressive performance. However, previous pre-training methods either oversimplified structured data into a sequence without considering input structures or designed training objectives tailored for a specific data structure (e.g., table or knowledge graph). In this paper, we unify different types of structured data (i.e., table, key-value data, knowledge graph) into the graph format and cast different D2T generation tasks as graph-to-text generation. To effectively exploit the structural information of the input graph, we propose a structure-enhanced pre-training method for D2T generation by designing a structure-enhanced Transformer. Concretely, we devise a position matrix for the Transformer, encoding relative positional information of connected nodes in the input graph. In addition, we propose a new attention matrix to incorporate graph structures into the original Transformer by taking the available explicit connectivity structure into account. Extensive experiments on six benchmark datasets show the effectiveness of our model. Our source codes are available at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/unid2t. Shujie Li 0001, Liang Li 0006, Ruiying Geng, Min Yang 0007, Binhua Li, Guanghu Yuan, Wanwei He, Shao Yuan, Can Ma, Fei Huang 0002, Yongbin Li 0001 |
Trans. Assoc. Comput. Linguistics | 9 |
| 2024 | A Survey on Model Compression for Large Language ModelsabstractAbstract Large Language Models (LLMs) have transformed natural language processing tasks successfully. Yet, their large size and high computational needs pose challenges for practical use, especially in resource-limited settings. Model compression has emerged as a key research area to address these challenges. This paper presents a survey of model compression techniques for LLMs. We cover methods like quantization, pruning, and knowledge distillation, highlighting recent advancements. We also discuss benchmarking strategies and evaluation metrics crucial for assessing compressed LLMs. This survey offers valuable insights for researchers and practitioners, aiming to enhance efficiency and real-world applicability of LLMs while laying a foundation for future advancements. Xunyu Zhu, Can Ma |
Trans. Assoc. Comput. Linguistics | 4 |
| 2023 | CATS: A Pragmatic Chinese Answer-to-Sequence Dataset with Large Scale and High QualityabstractLiang Li, Ruiying Geng, Chengyang Fang, Bing Li, Can Ma, Rongyu Cao, Binhua Li, Fei Huang, Yongbin Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liang Li 0006, Ruiying Geng, Chengyang Fang, Bing Li 0001, Can Ma, Rongyu Cao, Binhua Li, Fei Huang 0002, Yongbin Li 0001 |
ACL (1) | 5 |
| 2023 | CACluster: A Clustering Approach for IoT Attack Activities Based on Contextual AnalysisabstractAttacks against IoT have shown a rapid increase in both quantity and complexity. Analysts must handle massive alerts and determine the type of attack manually. In addition, the same attack activity may present polymorphism alert sequences due to overlapping attacks, adaptive attack strategy, error alerts, etc, which poses a severe challenge for human analysis. This manual-dependent and scenario-by-scenario security model is seriously overwhelming security analysts. This paper proposes a contextual-analysis-based clustering approach, CACluster, to aggregate similar attack activities end-to-end. It embeds alert context into vector space and uses an unsupervised clustering method to find similar attack activities based on domain matching and vector distance. Experimental results demonstrate that the CACluster could accurately aggregate similar attack activities, with 0.888 purity, reducing the number of attack activities by 84.8%. It will significantly cut down analysts’ workload. Huiran Yang, Yan Zhang 0014, Yueyue Dai, Jiyan Sun, Huajun Cui, Can Ma, Weiping Wang 0005 |
ICPADS | 6 |
| 2023 | An Unsupervised Vision-related Keywords Retrieval and Fusion Method for Visual StorytellingabstractVisual storytelling is a multi-modal generation task aiming to generate a coherent story for a sequence of images. Previous visual storytelling models utilize task-beneficial non-visual features, e.g., emotion, sentiment, or knowledge graph, as additional supplements for visual features to further improve the generation quality of stories. However, these appropriate non-visual features require to be carefully designed and selected by specialized researchers, moreover, high-quality external knowledge sources are not readily available. This increases the development cost of the VST model. To alleviate the above problem, this paper explores mining the knowledge existing in the multi-modal pre-trained models (MM-PTMs). First, we propose an Unsupervised Keywords Retrieval module (UKR), which takes an MM-PTM as the expert to select image-related keywords from a prepared task-related text corpus. The retrieved keywords not only complement and illustrate the visual features of the images but also provide more explicit generative signals to improve the interpretability and controllability of the generation process. Furthermore, we propose a Local Multi-modal Adaptive Fusion module (LMAF) to better fuse the textual and visual features and avoid noise brought by irrelevant keywords. LMAF dynamically aggregates features from both modalities through finer-grained correlation matching. The experimental results on the VST dataset, VIST, show that our proposed method achieves competitive results on several automatic metrics. Comparable results to previous methods can be achieved even if the model does not refer to visual features during story generation. Can Ma, Xiyan Gao, Guangheng Jia |
ICTAI | 2 |
| 2023 | Separate and Locate: Rethink the Text in Text-based Visual Question AnsweringabstractText-based Visual Question Answering (TextVQA) aims at answering questions about the text in images. Most works in this field focus on designing network structures or pre-training tasks. All these methods list the OCR texts in reading order (from left to right and top to bottom) to form a sequence, which is treated as a natural language ''sentence''. However, they ignore the fact that most OCR words in the TextVQA task do not have a semantical contextual relationship. In addition, these approaches use 1-D position embedding to construct the spatial relation between OCR tokens sequentially, which is not reasonable. The 1-D position embedding can only represent the left-right sequence relationship between words in a sentence, but not the complex spatial position relationship. To tackle these problems, we propose a novel method named Separate and Locate (SaL) that explores text contextual cues and designs spatial position embedding to construct spatial relations between OCR texts. Specifically, we propose a Text Semantic Separate (TSS) module that helps the model recognize whether words have semantic contextual relations. Then, we introduce a Spatial Circle Position (SCP) module that helps the model better construct and reason the spatial position relationships between OCR texts. Our SaL model outperforms the baseline model by 4.44% and 3.96% accuracy on TextVQA and ST-VQA datasets. Compared with the pre-training state-of-the-art method pre-trained on 64 million pre-training samples, our method, without any pre-training tasks, still achieves 2.68% and 2.52% accuracy improvement on TextVQA and ST-VQA. Our code and models will be released at https://github.com/fangbufang/SaL. Chengyang Fang, Can Ma, Dayong Hu |
ACM Multimedia | 4 |
| 2023 | Feature Enhancement with Text-Specific Region Contrast for Scene Text Detection
Xurui Sun, Jiahao Lyu 0002, Yifei Zhang 0005, Gangyan Zeng, Bo Fang 0003, Yu Zhou 0015, Enze Xie, Can Ma |
PRCV (7) | 8 |
| 2023 | LActDet: An Automatic Network Attack Activity Detection Framework for Multi-step AttacksabstractWith the evolution of attack tactics, cyber-attacks are presenting a sophisticated trend. The multi-step attack has become the mainstream attack form, where adversaries implement multiple attack steps to achieve their goals, which poses server challenges to attack detection. Traditional research mainly concentrates on how a particular attack step is exploited but fails to identify the whole attack activity automatically. Manual analysis is required to correlate multiple steps and determine the fine-grained type of attack activities, which is a heavy workload. In addition, the high error rate of alerts results in a negative impact on attack-activity detection performance.To address these challenges, we propose a framework, LActDet, to automatically identify attack activities from the raw alerts end-to-end. Firstly, it utilizes a document-embedding method to vectorize attack-event descriptions. Second, a seq2seq model is implemented to embed the attack-event sequence into the attack-phase sequence to represent the framework of attack activity, aiming at improving the fault tolerance for error alerts. In the end, we propose a temporal-sequence-based classifier to identify attack activities. Our experimental results demonstrate that LActDet achieves higher detection accuracy, lower artificial dependence, and less system overhead. Huiran Yang, Jiaqi Kang, Yueyue Dai, Jiyan Sun, Yan Zhang 0014, Huajun Cui, Can Ma |
TrustCom | 7 |
| 2022 | Graph-to-Text Generation with Dynamic Structure PruningabstractMost graph-to-text works are built on the encoder-decoder framework with cross-attention mechanism. Recent studies have shown that explicitly modeling the input graph structure can significantly improve the performance. However, the vanilla structural encoder cannot capture all specialized information in a single forward pass for all decoding steps, resulting in inaccurate semantic representations. Meanwhile, the input graph is flatted as an unordered sequence in the cross attention, ignoring the original graph structure. As a result, the obtained input graph context vector in the decoder may be flawed. To address these issues, we propose a Structure-Aware Cross-Attention (SACA) mechanism to re-encode the input graph representation conditioning on the newly generated context at each decoding step in a structure aware manner. We further adapt SACA and introduce its variant Dynamic Graph Pruning (DGP) mechanism to dynamically drop irrelevant nodes in the decoding process. We achieve new state-of-the-art results on two graph-to-text datasets, LDC2020T02 and ENT-DESC, with only minor increase on computational cost. Ruiying Geng, Can Ma, Yinliang Yue, Binhua Li |
COLING | 4 |
| 2022 | Relation-Aware Global-Augmented Transformer for TextCaps
Can Ma |
ICANN (1) | 3 |
| 2022 | Towards Escaping from Language Bias and OCR Error: Semantics-Centered Text Visual Question AnsweringabstractTexts in scene images convey critical information for scene understanding and reasoning. The abilities of reading and rea-soning matter for the model in the text-based visual question answering (TextVQA) process. However, current TextVQA models do not center on the text and suffer from several limitations. The model is easily dominated by language biases and optical character recognition (OCR) errors due to the ab-sence of semantic guidance in the answer prediction process. In this paper, we propose a novel Semantics-Centered Net-work (SC-Net) that consists of an instance-level contrastive semantic prediction module (ICSP) and a semantics-centered transformer module (SCT). Equipped with the two modules, the semantics-centered model can resist the language biases and the accumulated errors from OCR. Extensive experiments on TextVQA and ST-VQA datasets show the effectiveness of our model. SC- Net surpasses previous works with a notice-able margin and is more reasonable for the TextVQA task. Chengyang Fang, Gangyan Zeng, Yu Zhou 0015, Daiqing Wu, Can Ma, Dayong Hu, Weiping Wang 0005 |
ICME | 5 |
| 2021 | Improving Encoder by Auxiliary Supervision Tasks for Table-to-Text GenerationabstractLiang Li, Can Ma, Yinliang Yue, Dayong Hu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Can Ma, Yinliang Yue, Dayong Hu |
ACL/IJCNLP (1) | 2 |
| 2021 | HSAN: A Hierarchical Self-Attention Network for Multi-Turn Dialogue GenerationabstractIn the multi-turn dialogue system, response generation is not only related to the sentences in context but also relies on the words in each utterance. Although there are lots of methods that pay attention to model words and utterances, there still exist problems such as tending to generate common responses. In this paper, we propose a hierarchical self-attention network, named HSAN, which attends to the important words and utterances in context simultaneously. Firstly, we use the hierarchical encoder to update the word and utterance representations with their position information respectively. Secondly, the response representations are updated by the mask self-attention module in the decoder. Finally, the relevance between utterances and response is computed by another self-attention module and used for the next response decoding process. In terms of automatic metrics and human judgements, experimental results show that HSAN significantly outperforms all baselines on two common public datasets. Yawei Kong, Lu Zhang 0084, Can Ma, Cong Cao 0001 |
ICASSP | 3 |
| 2021 | CSPN: Multi-Scale Cascade Spatial Pyramid Network for Object DetectionabstractScale variation is one of the key challenges in object detection. One solution is Image Pyramid, which employs images of multiple resolutions for training. Another solution is Feature Pyramid, which uses multi-scale features for prediction and is widely used in current object detectors due to its high efficiency. However, the representational power of each scale in Feature Pyramid is inconsistent, which makes the performance lower than Image Pyramid. To solve this problem and obtain better detection performance, we propose a novel net-work named Multi-Scale Cascade Spatial Pyramid Network (MS-CSPN) to strengthen Feature Pyramid. First, we de-sign CSPN to expand the receptive field in a cascade form to detect objects of different scales. Secondly, we propose a Cross-Scale Sharing Strategy, which shares the parameters of CSPN at all scales. Finally, we introduce global context information to enhance MS-CSPN. Experimental results on the MS-COCO benchmark show that the proposed MS-CSPN improves the mAP by a large margin compared to previous related works. Tianyuan Wang, Can Ma, Haoshan Su, Weiping Wang 0005 |
ICASSP | 2 |
| 2021 | SSFENet: Spatial and Semantic Feature Enhancement Network for Object DetectionabstractCurrent state-of-the-art object detectors generally use pre-trained classification networks to extract features, and then utilize feature pyramids to detect objects of different scales. However, classification networks prefer translation invariance and ignore the location information, so directly using the extracted features for fusion will affect the performance. In this paper, we present a novel network to address this dilemma, denoted as Spatial and Semantic Feature Enhancement Network (SSFENet). First, we introduce Spatial Feature Enhancement Block to utilize dilated convolution and weighted feature fusion to enhance the spatial information in features. Second, in the low-level stage, our Semantic Feature Enhancement Block uses the backbone network of the high-level stage to obtain features with richer semantic information and only introduces little computational cost due to the use of shared convolution layers. Experimental results on the MS-COCO benchmark show that the proposed SSFENet significantly improves the mAP of commonly used object detectors. Tianyuan Wang, Can Ma, Haoshan Su, Weiping Wang 0005 |
ICASSP | 2 |
| 2020 | Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningabstractWe propose a novel self-supervised method, referred to as Video Cloze Procedure (VCP), to learn rich spatial-temporal representations. VCP first generates “blanks” by withholding video clips and then creates “options” by applying spatio-temporal operations on the withheld clips. Finally, it fills the blanks with “options” and learns representations by predicting the categories of operations applied on the clips. VCP can act as either a proxy task or a target task in self-supervised learning. As a proxy task, it converts rich self-supervised representations into video clip operations (options), which enhances the flexibility and reduces the complexity of representation learning. As a target task, it can assess learned representation models in a uniform and interpretable manner. With VCP, we train spatial-temporal representation models (3D-CNNs) and apply such models on action recognition and video retrieval tasks. Experiments on commonly used benchmarks show that the trained models outperform the state-of-the-art self-supervised models with significant margins. Dezhao Luo, Chang Liu 0042, Yu Zhou 0015, Dongbao Yang, Can Ma, Qixiang Ye, Weiping Wang 0005 |
AAAI | 5 |
| 2020 | Detecting Manipulated Facial Videos: A Time Series SolutionabstractWe propose a new method to expose fake videos based on a time series solution. The method is based on bidirectional long short-term memory (Bi-LSTM) backbone architecture with two different types of features: Face-Alignment and Dense-Face-Alignment, in which both of them are physiological signals that can be distinguished between fake and original videos. We choose 68 landmark points as the feature of Face-Alignment and Pose Adaptive Feature (PAF) for Dense-Face-Alignment. Based on these two facial features, we designed two deep networks. In addition, we optimize our network by adding an attention mechanism that improves detection precision. Our method is tested over benchmarks of Face Forensics/Face Forensics++ dataset and show a promising performance on inference speed while maintaining accuracy with state-of art solutions that deal against DeepFake. Can Ma, Meilin Gao |
ICPR | 2 |
| 2020 | From When to Where: A Multi-task Learning Approach for Next Point-of-Interest Recommendation
Jinwen Zhong, Can Ma, Weiping Wang 0005 |
WASA (1) | 2 |
| 2015 | An approach of fast data manipulation in HDFS with supplementary mechanisms
Can Ma, Weiping Wang 0005, Dan Meng 0002 |
J. Supercomput. | 2 |
| 2013 | Zput: A speedy data uploading approach for the Hadoop Distributed File SystemabstractHadoop Distributed File System (HDFS) is the storage component of the Hadoop framework, which is designed for maintaining and processing huge datasets efficiently among cluster nodes. To cooperate with MapReduce, the computation infrastructure of Hadoop, data is required to be uploaded from local file systems to HDFS. Unfortunately when data is of massive scale, the uploading procedure becomes extremely time-consuming, which causes serious delay for urgent tasks. This primary contribution of this paper is the proposition of Zput, a speedy data uploading mechanism which can significantly accelerate uploading by using metadata mapping approach. After the implementation is described and corresponding advantages are narrated, disadvantages are also analyzed and eliminated by using an approach named remote block placement. Evaluation results show this new mechanism can reduce the running time of uploading process by about 60-90%, and the remote block placement can boost the course of block distribution by about 30-40%, while maintaining the complete compatibility for upper-layer applications. Weiping Wang 0005, Can Ma, Dan Meng 0002 |
CLUSTER | 3 |
| 2013 | An overlapping clustering approach for routing in Wireless Sensor NetworksabstractThe design and analysis of routing algorithm is an important issue in Wireless Sensor Networks (WSNs). Most traditional geographical routing algorithms cannot achieve good performance in duty-cycled networks. In this paper, we propose a k-connected overlapping clustering approach with energy awareness, namely k-OCHE, for routing in WSNs. The basic idea of this approach is to select a cluster head by energy availability (EA) status. The k-OCHE scheme adopts a sleep scheduling strategy of CKN, where neighbors will remain awake to keep it k-connected, so that it can balance energy distributions well. Compared with traditional routing algorithms, the proposed k-OCHE approach obtains a balanced load distribution and consequently a longer network lifetime. Can Ma, Lei Wang 0005, Zhenquan Qin, Lei Shu 0001, Di Wu 0007 |
WCNC | 1 |
| 2012 | Clover: A Distributed File System of Expandable Metadata Service Derived from HDFSabstractTo store and manage data efficiently is the critical issue which modern information infrastructures confront with. To accommodate the massive scale of data in the Internet environment, most common solutions utilize distributed file systems. However there still exist disadvantages preventing these systems from delivering satisfying performance. In this paper, we present a Name Node cluster file system based on HDFS, which is named Clover. This file system exploits two critical features: an improved 2PC protocol which ensures consistent metadata update on multiple metadata servers and a shared storage pool which provides robust persistent metadata storage and supports the operation of distributed transactions. Clover is compared with HDFS and its key virtues are shown. Further experimental results show our system can achieve better metadata expandability ranging from 10% to 90% by quantized metrics when each extra server is added, while preserving similar I/O performance. Can Ma, Weiping Wang 0005, Dan Meng 0002, Jason Kei |
CLUSTER | 3 |
| 2011 | A Geographic Routing Algorithm in Duty-Cycled Sensor Networks with Mobile SinksabstractIn this paper, we focus on achieving better energy conservation for geographic routing algorithms in duty-cycled WSNs when there is a mobile sink. We simplify the problem as a topology coverage one, and propose a multi-metric geographic algorithm (MMGR) which uses multi-metric candidates (MMCs) for geographic routing. The analysis and extensive simulation results show that MMGR can achieve better energy conservations than McTPGF, while retaining good performance of end-to-end delay and hop counts. Can Ma, Lei Wang 0005, Zhenquan Qin, Ming Zhu 0001, Lei Shu 0001 |
MSN | 1 |
| 2011 | HR-NET: A Highly Reliable Message-Passing Mechanism for Cluster File SystemabstractAs PC clusters increase in popularity and quantity, message-passing between nodes has been an important issue for high failure rate in the network. File access in a cluster file system often contains several sub-operations, each includes one or more network transmissions. Any network failures will cause the file system service unavailable. In this paper, we describe a highly reliable message-passing mechanism (HRNET), which tolerates both software and hardware network failures. HR-NET provides fine-grained, connection-level fail over across communication path redundancy. With it the file system can keep passing messages until it either recovers from network failures or it is failed over to a backup. Load balance for messages is also achieved to relieve network traffic. For transmission timeout, HR-NET proposes the message priority scheduling which dynamically manages messages in an appropriate order to tolerate request-response failures between clients and servers. As HR-NET is completely independent, there are neither any changes to standard protocol stacks nor modifications at upper file system. Performance results show that HR-NET takes full advantage of network bandwidth with average 6.17% throughput loss and provides a fast recovery. Experiments with cluster file system dispose that the overall performance degradation is below 8% due to failover of HR-NET while the reliability is highly enhanced. Can Ma, Jin Xiong |
NAS | 2 |
| 2011 | Dawning Nebulae: A PetaFLOPS Supercomputer with a Heterogeneous Structure
Ninghui Sun, Zhigang Huo, Guangming Tan, Jin Xiong, Bo Li 0009, Can Ma |
J. Comput. Sci. Technol. | 7 |
| 2009 | DCR: A fully transparent checkpoint/restart framework for distributed systemsabstractCheckpoint/restart has been widely used in computing systems for fault tolerance, job scheduling and system maintenance purposes. However, the lack of transparency has hindered adoptions of many implementations of it. In this paper, we present a fully transparent parallel checkpoint/restart framework, DCR, which takes the advantages of kernel-level checkpointing method and TCP session preservation. DCR is fully transparent to application programmers and users. No source code modifications, recompilations, or system call interceptions are required. Because of the simplicity of its design and the dominance of TCP/IP in parallel applications, DCR can be readily deployed in widely scales of computers, from single CPU computers to large-scale clusters. A new on-demand blocking checkpoint protocol, which makes use of the reliability mechanism of TCP, is proposed to eliminate the global synchronization. We have demonstrated the effectiveness and efficiency of DCR by multiple MPICH2 applications running on Dawning 5000A. Can Ma, Zhigang Huo, Jingnan Cai |
CLUSTER | 1 |
| 2008 | HPPNET: A novel network for HPC and its implication for communication softwareabstractWith the widespread adoption of multicore processors in high performance computing (HPC) environment, the balance between computation and communication moves towards computation. It has been becoming more important to design a high efficient network for HPC system, which commonly has two challengers associated: 1) To provide a communication environment with low latency, high bandwidth, and high small message processing rate; 2) and to efficiently support the partitioned global address space (PGAS) programming model. With respects to these needs this paper proposes a novel network named HPPNET. By adopting HyperTransport interface, separate channel design, on-load part of processing work to host, and transparent direct load/store in global physical address space, HPPNET can sufficiently support both needs for HPC. Meanwhile, we have adopted several key technologies to minimize the implication of new network for communication software. Evaluation shows that HPPNET hardware design can achieve high performance and bring no barrier to high efficiency software design. Our results also show that the 8 bytes remote store cost 0.4s in HPPNET prototype. Panyong Zhang, Can Ma |
IPDPS | 2 |