EDBT 2026 Demo / reviewers in the wild / expert
Shan Zhao 0002
dblp:00/6640-2
· DBLP profile ↗
28ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0003-4503-8259ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Image as Compressed Visual Prompt and Hierarchical Visual Knowledge for Effective Image Utilization in MLLMsabstractMultimodal Large Language Models (MLLMs) integrate text and images for complex reasoning tasks, but efficiently utilizing image remains a challenge due to redundancy and noise. Traditional methods take the entire image features as visual prompt into the MLLMs, leading to excessive visual tokens that disrupt textual information expression. Thus, recent studies treat image features as visual knowledge, storing them in the feed-forward network for retrieval when needed. These methods, completely removing images from the input, may hinder the activation of image-related knowledge. Besides, current visual knowledge focuses on fine-grained details but overlooks the hierarchical process of visual perception. As described in feature integration theory, global structure is first processed before details are integrated. Ignoring this process may lead to a fragmented visual understanding, making it difficult to capture high-level semantic relationships. To overcome these issues, we propose a novel image utilization mechanism in MLLMs. We leverage a compression-based attention mechanism to generate the compressed visual prompt, which not only mitigates the interference of excessively long visual prompts but also preserves crucial visual information necessary for activating knowledge in the MLLM. Furthermore, we extract hierarchical visual features as visual knowledge using wavelet transforms, allowing the model to capture both global structures and fine-grained details. Experiments show that our method achieves state-of-the-art performance. Shezheng Song, Kangcheng Ding, Shan Zhao 0002, Shasha Li 0001, Xiaopeng Li 0006, Chengyu Wang 0008, Qian Wan 0007, Bin Ji 0002, Jie Yu 0008 |
AAAI | 3 |
| 2026 | RICA: Re-ranking with intra-modal and cross-modal alignment for text-based person search
Yu Bai 0022, Wentao Ma 0003, Shan Zhao 0002, Tianwei Yan 0001, Shezheng Song, Chengyu Wang 0008, Qian Wan 0007 |
Expert Syst. Appl. | 3 |
| 2025 | Psychologically-Aware Retrieval-Augmented Generation for Coherent Role-Playing in LLMs
Pengyang Shao, Shan Zhao 0002, Shezheng Song, Tianwei Yan 0001, Chengyu Wang 0008 |
PRCV (4) | 3 |
| 2025 | Exploring Identifiable Tokens for Text-Based Person Re-identification
Wentao Ma 0003, Pengfei Yuan, Feng Li 0037, Shan Zhao 0002 |
PRCV (18) | 6 |
| 2025 | Hierarchical knowledge-guided reasoning for text-based person re-identification
Ruigeng Zeng, Wentao Ma 0003, Tongqing Zhou, Shan Zhao 0002, Xinjun Mao, Jie Liu 0002 |
Neural Networks | 4 |
| 2025 | ISSD: Indicator Selection for Time Series State DetectionabstractTime series data from monitoring applications captures the behaviours of objects, which can often be split into distinguishable segments that reflect the underlying state changes. Despite the recent advances in time series state detection, the indicator selection for state detection is rarely studied, most of state detection work assumes the input indicators have been properly or manually selected. However, this assumption is disconnected from practice, on one hand, manual selection is not scalable, there can be up to thousands of indicators for the runtime monitoring of certain objects, e.g., supercomputer systems. On the other hand, performing state detection on a large amount of raw indicators is both inefficient and redundant. We argue that indicator selection should be made an upstream task for selecting a subset of indicators to facilitate state analysis. To this end, we propose ISSD ( I ndicator S election for S tate D etection), an indicator selection method for time series state detection. At its core, ISSD attempts to find an indicator subset that has as much high-quality states, which is measured by the channel set completeness and quality we invent based on segment-level sampling statistics. Such an indicator selection process is transformed into a multi-objective optimization problem and an approximation algorithm is designed to solve the NP-hard searching for specific end point in the Pareto front. Experiments on 5 datasets and 4 downstream methods show that ISSD has significant selection superiority compared with 6 baselines. We also elaborate on two observations of selection resilience and channel sensitivity of existing state detection methods and appeal to further research on them. Chengyu Wang 0008, Tongqing Zhou, Lin Chen 0028, Shan Zhao 0002, Zhiping Cai |
Proc. ACM Manag. Data | 4 |
| 2025 | How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language ModelabstractWe explore Multimodal Large Language Models (MLLMs), which integrate LLMs like GPT-4 to handle multimodal data, including text, images, audio, and more. MLLMs demonstrate capabilities such as generating image captions and answering image-based questions, bridging the gap towards real-world human-computer interactions and hinting at a potential pathway to artificial general intelligence. However, MLLMs still face challenges in addressing the semantic gap in multimodal data, which may lead to erroneous outputs, posing potential risks to society. Selecting the appropriate modality alignment method is crucial, as improper methods might require more parameters without significant performance improvements. This paper aims to explore modality alignment methods for LLMs and their current capabilities. Implementing effective modality alignment can help LLMs address environmental issues and enhance accessibility. The study surveys existing modality alignment methods for MLLMs, categorizing them into four groups: (1) Multimodal Converter, which transforms data into a format that LLMs can understand; (2) Multimodal Perceiver, which improves how LLMs percieve different types of data; (3) Tool Learning, which leverages external tools to convert data into a common format, usually text; and (4) Data-Driven Method, which teaches LLMs to understand specific data types within datasets. Shezheng Song, Xiaopeng Li 0006, Shasha Li 0001, Shan Zhao 0002, Jie Yu 0008, Jun Ma 0015, Xiaoguang Mao, Meng Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Dual-Decoupling With Frequency-Spatial Domains for Image Manipulation LocalizationabstractLeveraging trace-rich features within embedded spaces has been established as effective in image manipulation localization (IML). Nevertheless, the feature of manipulated traces frequently comprises substantial redundant information only loosely related to IML tasks. This complexity has hindered existing methods in fully comprehending the essence of trace features. In light of this challenge, we introduce a novel decoupling representation learning network (DRN) tailored for IML. The DRN excels at decoupling intricate multidomain information and transforming it into representations directly pertinent to IML objectives. This is achieved through a meticulously designed frequency decoupling representation learning module (FDM) and spatial decoupling representation learning module (SDM). Specifically, the FDM operates by acquiring distinct low and high-frequency components to effectively decouple redundant information. The decoupled high-frequency components are then harnessed as intricate trace complements, enhancing the overall aggregation process. In addition, the redundant information is expertly separated into authentic and manipulated representations through the use of channel activation maps in SDM. Through extensive experimentation on three public benchmarks including CASIA, NIST, and Coverage, our method consistently demonstrates superior performance and enhanced robustness compared with existing state-of-the-art methods. Wenyan Pan, Wentao Ma 0003, Tongqing Zhou, Shan Zhao 0002, Lichuan Gu, Guolong Shi, Zhihua Xia |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Hierarchical Label-Enhanced Contrastive Learning for Chinese NERabstractRecently, character-word lattice structures have achieved promising results for Chinese named entity recognition (NER), reducing word segmentation errors and increasing word boundary information for character sequences. However, constructing the lattice structure is complex and time-consuming, thus these lattice-based models usually suffer from low inference speed. Moreover, the quality of the lexicon affects the accuracy of the NER model. Since noise words can potentially confuse NER, limited coverage of the lexicon can cause lattice-based models to degenerate into partial character-based models. In this article, we propose a hierarchical label-enhanced contrastive learning (HLCL) method for Chinese NER. Instead of relying on the lattice structure, HLCL offers an alternative solution to robustly integrate entity boundary and type information with the help of both labels semantic and contrastive learning. HLCL is empowered by two techniques: 1) sentence-level contrastive learning (SCL) to model global mutual information between two different modalities (e.g., labels and sentences) and 2) token-level contrastive learning (TCL) to close the gap between representations of different characters (e.g., label-enhanced characters and original characters), resulting in local mutual information. With the well-designed contrastive learning scheme and the concise model during inference, HLCL can fully leverage the transferable label semantic and has a superb speed of inference. Experiments on four Chinese NER datasets show that HLCL obtains excellent efficiency as well as performance compared with existing lattice-based approaches. Chengyu Wang 0008, Shan Zhao 0002, Tianwei Yan 0001, Shezheng Song, Wentao Ma 0003, Kuien Liu, Meng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | HCL: A Hierarchical Contrastive Learning Framework for Zero-Shot Relation ExtractionabstractZero-shot relation extraction (ZSRE) is shown to become more significant in the current information extraction system, which aims at predicting relation classes that lack annotations or have just never appeared during training. Previous works focus on projecting sentences with their corresponding relation descriptions to an intermediate semantic space and searching the nearest semantic for predicting unseen classes. Though these methods can achieve sound performance, they only obtain inferior semantic information via a trivial distance metric and neglect the interaction in the instance representations. We are thus motivated to tackle these issues and propose a hierarchical contrastive learning (HCL) framework for ZSRE including projection-level and instance-level modules. Specifically, the projection-level component replaces the distance score function by contrastive loss to connect the input sentence with the relation semantic space. And the instance-level component integrates the external knowledge from sentence entities to establish new contrastive pairs for efficiently learning representations from mutual information. The experimental results on three well-known datasets demonstrate that our model surpasses the existing SOTA by at most 18.97% improvement on the F1 score when unseen classes are 15. Moreover, our model can achieve more competitive performance alone with the increasing number of unseen classes. Tianwei Yan 0001, Shan Zhao 0002, Minghao Hu 0001, Mengzhu Wang, Xiang Zhang 0008, Zhigang Luo, Meng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | FRCL-MNER: A Finer Grained Rank-Based Contrastive Learning Framework for Multimodal NERabstractMultimodal named entity recognition (MNER) is an emerging field that aims to automatically detect named entities and classify their categories, utilizing input text and auxiliary resources such as images. While previous studies have leveraged object detectors to preprocess images and fuse textual semantics with corresponding image features, these methods often overlook the potential finer grained information within each modality and may exacerbate error propagation due to predetection. To address these issues, we propose a finer grained rank-based contrastive learning (FRCL) framework for MNER. This framework employs a global-level contrastive learning to align multimodal semantic features and a Top-K rank-based mask strategy to construct positive-negative pairs, thereby learning a finer grained multimodal interaction representation. Experimental results from three well-known social media datasets reveal that our approach surpasses existing strong baselines, and achieves up to a 1.54% improvement on the Twitter2015 dataset. Extensive discussions further confirm the effectiveness of our approach. We will release the source code on https://github.com/augusyan/FRCL. Tianwei Yan 0001, Shan Zhao 0002, Wentao Ma 0003, Shezheng Song, Chengyu Wang 0008, Zhibo Rao, Shizhao Chen, Zhigang Luo, Xinwang Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity LinkingabstractMultimodal Entity Linking (MEL) aims at linking ambiguous mentions with multimodal information to entity in Knowledge Graph (KG) such as Wikipedia, which plays a key role in many applications. However, existing methods suffer from shortcomings, including modality impurity such as noise in raw image and ambiguous textual entity representation, which puts obstacles to MEL. We formulate multimodal entity linking as a neural text matching problem where each multimodal information (text and image) is treated as a query, and the model learns the mapping from each query to the relevant entity from candidate entities. This paper introduces a dual-way enhanced (DWE) framework for MEL: (1) our model refines queries with multimodal data and addresses semantic gaps using cross-modal enhancers between text and image information. Besides, DWE innovatively leverages fine-grained image attributes, including facial characteristic and scene feature, to enhance and refine visual features. (2)By using Wikipedia descriptions, DWE enriches entity semantics and obtains more comprehensive textual representation, which reduces between textual representation and the entities in KG. Extensive experiments on three public benchmarks demonstrate that our method achieves state-of-the-art (SOTA) performance, indicating the superiority of our model. The code is released on https://github.com/season1blue/DWE. Shezheng Song, Shan Zhao 0002, Chengyu Wang 0008, Tianwei Yan 0001, Shasha Li 0001, Xiaoguang Mao, Meng Wang 0001 |
AAAI | 2 |
| 2024 | Text-Based Occluded Person Re-identification via Multi-Granularity Contrastive Consistency LearningabstractText-based Person Re-identification (T-ReID), which aims at retrieving a specific pedestrian image from a collection of images via text-based information, has received significant attention. However, previous research has overlooked a challenging yet practical form of T-ReID: dealing with image galleries mixed with occluded and inconsistent personal visuals, instead of ideal visuals with a full-body and clear view. Its major challenges lay in the insufficiency of benchmark datasets and the enlarged semantic gap incurred by arbitrary occlusions and modality gap between text description and visual representation of the target person. To alleviate these issues, we first design an Occlusion Generator (OGor) for the automatic generation of artificial occluded images from generic surveillance images. Then, a fine-granularity token selection mechanism is proposed to minimize the negative impact of occlusion for robust feature learning, and a novel multi-granularity contrastive consistency alignment framework is designed to leverage intra-/inter-granularity of visual-text representations for semantic alignment of occluded visuals and query texts. Experimental results demonstrate that our method exhibits superior performance. We believe this work could inspire the community to investigate more dedicated designs for implementing T-ReID in real-world scenarios. The source code is available at https://github.com/littlexinyi/MGCC. Wentao Ma 0003, Dan Guo 0001, Tongqing Zhou, Shan Zhao 0002, Zhiping Cai |
AAAI | 5 |
| 2024 | DIM: Dynamic Integration of Multimodal Entity Linking with Large Language Model
Shezheng Song, Shasha Li 0001, Jie Yu 0008, Shan Zhao 0002, Xiaopeng Li 0006, Jun Ma 0015, Xiaodong Liu 0004, Xiaoguang Mao |
PRCV (5) | 4 |
| 2024 | Auto-focus tracing: Image manipulation detection with artifact graph contrastive
Wenyan Pan, Zhihua Xia, Wentao Ma 0003, Yuwei Wang 0002, Lichuan Gu, Guolong Shi, Shan Zhao 0002 |
Knowl. Based Syst. | 7 |
| 2024 | Document-level relation extraction with three channels
Zhanjun Zhang, Shan Zhao 0002, Qian Wan 0007, Jie Liu 0002 |
Knowl. Based Syst. | 2 |
| 2024 | Image Manipulation Detection With Cascade Hierarchical Graph RepresentationabstractRecent image manipulation detection approaches primarily rely on sophisticated Convolutional Neural Network (CNN)-based models for region localization, while they tend to ignore: 1) the feature correlations that exist between manipulated and non-manipulated regions. 2) the significance of multi-scale representations in detecting manipulated regions of varying sizes, consequently hampering the overall performance of image manipulation detection. To address these limitations, we propose a novel approach, called Cascade Hierarchical Graph Convolutional Network (Cas-HGCN), which comprehensively learns the feature correlations between manipulated and non-manipulated regions at different scales using the Feature Correlations Modeling (FCM) module. Specifically, the FCM module treats the grids in the hierarchical image/feature maps as nodes, constructs a fully-connected graph by connecting each node, and leverages it to learn and refine feature correlations across different scales in a cascading manner. This process results in high discriminability for distinguishing manipulated and non-manipulated regions. Extensive experiments conducted on three public datasets, namely CASIA, NIST, and Coverage, demonstrate the promising detection accuracy achieved by Cas-HGCN without the need for pre-training on large datasets, surpassing the performance of existing state-of-the-art competitors. Wenyan Pan, Wentao Ma 0003, Shan Zhao 0002, Lichuan Gu, Guolong Shi, Zhihua Xia, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Remote Sensing Image-Text Retrieval With Implicit-Explicit Relation ReasoningabstractRemote sensing image-text retrieval (RSITR) has become a research hotspot in recent years for its wide application. Existing methods in this context, based either on local or global feature matching, overlook the sensing variation-leaded visual deviation and geographically nearby image-text mismatching problems of remote sensing (RS) images. This work notes that this would limit the retrieval accuracy for RSITR. To handle this, we present IERR, an implicit-explicit relation reasoning framework that learns relations between local visual-textual tokens and enhances global image-text matching without requiring additional prior supervision. Specifically, masked image modeling (MIM) and masked language modeling (MLM) are used for symmetric mask reasoning consistency alignment. Meanwhile, masked features (i.e., implicit relation) and unmasked features (i.e., explicit relation) are fed into a multimodal interaction encoder to enhance the representations of the textual-visual features. Extensive experimental results on the RSICD and RSITMD datasets demonstrate the superiority of IERR compared with 17 baselines. Lingling Yang, Tongqing Zhou, Wentao Ma 0003, Mengze Du, Lu Liu 0023, Feng Li 0037, Shan Zhao 0002, Yuwei Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | FedSH: Towards Privacy-Preserving Text-Based Person Re-IdentificationabstractText-based person re-identification (ReID) has enabled canonical applications in searching for and tracking targets from large-scale surveillance images with textual descriptions. Yet, existing text-based person ReID systems employ centralized model training that gathers images captured by different institutes' cameras into one place, which poses severe privacy threats to sensitive institutional information. This work is then devoted to exploring privacy-preserving text-based person ReID and proposes the framework of FedSH by tailoring the federated learning paradigm for distributed searching knowledge extraction. Specifically, FedSH resolves the local model generalization and entity boundary obscuring limitations, caused by inner-institute data homogeneity and inter-institute data heterogeneity, via building multi-granularity feature representation and a semantically self-aligned network. Meanwhile, it reduces the communication burden introduced by the embedding for multiple modals by updating common representation subspaces during federated learning. Experimental results on two public benchmarks demonstrate that our method can achieve at most 16.47% and 16.02% person ReID performance improvement by the Rank-1 metric, compared with 6 State-of-The-Art (SoTA) baselines and 6 ablation studies. We believe that our work will inspire the community to investigate the potential of implementing Federated Learning in real-world image retrieval and ReID scenarios. Wentao Ma 0003, Shan Zhao 0002, Tongqing Zhou, Dan Guo 0001, Lichuan Gu, Zhiping Cai, Meng Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | MCL: Multi-Granularity Contrastive Learning Framework for Chinese NERabstractRecently, researchers have applied the word-character lattice framework to integrated word information, which has become very popular for Chinese named entity recognition (NER). However, prior approaches fuse word information by different variants of encoders such as Lattice LSTM or Flat-Lattice Transformer, but are still not data-efficient indeed to fully grasp the depth interaction of cross-granularity and important word information from the lexicon. In this paper, we go beyond the typical lattice structure and propose a novel Multi-Granularity Contrastive Learning framework (MCL), that aims to optimize the inter-granularity distribution distance and emphasize the critical matched words in the lexicon. By carefully combining cross-granularity contrastive learning and bi-granularity contrastive learning, the network can explicitly leverage lexicon information on the initial lattice structure, and further provide more dense interactions of across-granularity, thus significantly improving model performance. Experiments on four Chinese NER datasets show that MCL obtains state-of-the-art results while considering model efficiency. The source code of the proposed method is publicly available at https://github.com/zs50910/MCL Shan Zhao 0002, Chengyu Wang 0008, Minghao Hu 0001, Tianwei Yan 0001, Meng Wang 0001 |
AAAI | 1 |
| 2023 | A Span-based Multi-Modal Attention Network for joint entity-relation extraction
Qian Wan 0007, Luona Wei, Shan Zhao 0002, Jie Liu 0002 |
Knowl. Based Syst. | 3 |
| 2023 | Using Multimodal Contrastive Knowledge Distillation for Video-Text RetrievalabstractCross-modal retrieval aims to enable a flexible bi-directional retrieval experience across different modalities (e.g., searching for videos with texts). Many existing efforts tend to learn a common semantic representation embedding space in which items of different modalities can be directly compared, wherein the positive global representations of video-text pairs are pulled close while the negative ones are pushed apart via pair-wise ranking loss. However, such a vanilla loss would unfortunately yield ambiguous feature embeddings for texts of different videos, causing inaccurate cross-modal matching and unreliable retrievals. Toward this end, we propose a multimodal contrastive knowledge distillation method for instance video-text retrieval, called MCKD, by adaptively using the general knowledge of self-supervised model (teacher) to calibrate mixed boundaries. Specifically, the teacher model is tailored for robust (less-ambiguous) visual-text joint semantic space by maximizing mutual information of co-occurred modalities during multimodal contrastive learning. This robust and structural inter-instance knowledge is then distilled, with the help of explicit discrimination loss, to a student model for improved matching performance. Extensive experiments on four public benchmark video-text datasets (MSR-VTT, TGIF, VATEX, and Youtube2Text) demonstrate that our MCKD can achieve at most 8.8%, 6.4%, 5.9%, and 5.3% improvement in text-to-video performance by the$\text{R}\text{@}1$metric, compared with 14 SoTA baselines. Wentao Ma 0003, Qingchao Chen, Tongqing Zhou, Shan Zhao 0002, Zhiping Cai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Dynamic Modeling Cross-Modal Interactions in Two-Phase Prediction for Entity-Relation ExtractionabstractJoint extraction of entities and their relations benefits from the close interaction between named entities and their relation information. Therefore, how to effectively model such cross-modal interactions is critical for the final performance. Previous works have used simple methods, such as label-feature concatenation, to perform coarse-grained semantic fusion among cross-modal instances but fail to capture fine-grained correlations over token and label spaces, resulting in insufficient interactions. In this article, we propose a dynamic cross-modal attention network (CMAN) for joint entity and relation extraction. The network is carefully constructed by stacking multiple attention units in depth to dynamic model dense interactions over token-label spaces, in which two basic attention units and a novel two-phase prediction are proposed to explicitly capture fine-grained correlations across different modalities (e.g., token-to-token and label-to-token). Experiment results on the CoNLL04 dataset show that our model obtains state-of-the-art results by achieving 91.72% F1 on entity recognition and 73.46% F1 on relation classification. In the ADE and DREC datasets, our model surpasses existing approaches by more than 2.1% and 2.54% F1 on relation classification. Extensive analyses further confirm the effectiveness of our approach. Shan Zhao 0002, Minghao Hu 0001, Zhiping Cai, Fang Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Enhancing Chinese Character Representation With Lattice-Aligned AttentionabstractWord-character lattice models have been proved to be effective for some Chinese natural language processing (NLP) tasks, in which word boundary information is fused into character sequences. However, due to the inherently unidirectional sequential nature, prior approaches have only learned sequential interactions of character-word instances but fail to capture fine-grained correlations in word-character spaces. In this article, we propose a lattice-aligned attention network (LAN) that aims to model dense interactions over word-character lattice structure for enhancing character representations. By carefully combining cross-lattice module, gated word-character semantic fusion unit, and self-lattice attention module, the network can explicitly capture fine-grained correlations across different spaces (e.g., word-to-character and character-to-character), thus significantly improving model performance. Experimental results on three Chinese NLP benchmark tasks demonstrate that LAN obtains state-of-the-art results compared to several competitive approaches. Shan Zhao 0002, Minghao Hu 0001, Zhiping Cai, Zhanjun Zhang, Tongqing Zhou, Fang Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Dynamic Modeling Cross- and Self-Lattice Attention Network for Chinese NERabstractWord-character lattice models have been proved to be effective for Chinese named entity recognition (NER), in which word boundary information is fused into character sequences for enhancing character representations. However, prior approaches have only used simple methods such as feature concatenation or position encoding to integrate word-character lattice information, but fail to capture fine-grained correlations in word-character spaces. In this paper, we propose DCSAN, a Dynamic Cross- and Self-lattice Attention Network that aims to model dense interactions over word-character lattice structure for Chinese NER. By carefully combining cross-lattice and self-lattice attention modules with gated word-character semantic fusion unit, the network can explicitly capture fine-grained correlations across different spaces (e.g., word-to-character and character-to-character), thus significantly improving model performance. Experiments on four Chinese NER datasets show that DCSAN obtains stateof-the-art results as well as efficiency compared to several competitive approaches. Shan Zhao 0002, Minghao Hu 0001, Zhiping Cai, Haiwen Chen, Fang Liu 0002 |
AAAI | 1 |
| 2021 | Leveraging Multi-granularity Heterogeneous Graph for Chinese Electronic Medical Records Summarization
Rui Luo 0006, Ye Wang 0023, Shan Zhao 0002, Tongqing Zhou, Zhiping Cai |
ICONIP (5) | 3 |
| 2020 | Modeling Dense Cross-Modal Interactions for Joint Entity-Relation ExtractionabstractJoint extraction of entities and their relations benefits from the close interaction between named entities and their relation information. Therefore, how to effectively model such cross-modal interactions is critical for the final performance. Previous works have used simple methods such as label-feature concatenation to perform coarse-grained semantic fusion among cross-modal instances, but fail to capture fine-grained correlations over token and label spaces, resulting in insufficient interactions. In this paper, we propose a deep Cross-Modal Attention Network (CMAN) for joint entity and relation extraction. The network is carefully constructed by stacking multiple attention units in depth to fully model dense interactions over token-label spaces, in which two basic attention units are proposed to explicitly capture fine-grained correlations across different modalities (e.g., token-to-token and labelto-token). Experiment results on CoNLL04 dataset show that our model obtains state-of-the-art results by achieving 90.62% F1 on entity recognition and 72.97% F1 on relation classification. In ADE dataset, our model surpasses existing approaches by more than 1.9% F1 on relation classification. Extensive analyses further confirm the effectiveness of our approach. Shan Zhao 0002, Minghao Hu 0001, Zhiping Cai, Fang Liu 0002 |
IJCAI | 1 |
| 2019 | Adversarial training based lattice LSTM for Chinese clinical named entity recognition
Shan Zhao 0002, Zhiping Cai, Haiwen Chen, Ye Wang 0023, Fang Liu 0002, Anfeng Liu |
J. Biomed. Informatics | 1 |