VLDB 2026 Research / reviewers in the wild / expert
Zequn Zhang
dblp:120/9628
· DBLP profile ↗
75ranked-venue papers
3as first author
72since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 10 · 8 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Prompted, Zero-Annotation SAM: Weakly Supervised Binary Medical Image Segmentation
Jinliang Su, Zequn Zhang |
ICIC (1) | 3 |
| 2026 | Context-Preserving Dermoscopic Editing: Mask-Guided Lesion-Aware Diffusion for Attribute ModificationabstractPrecise manipulation of dermoscopic attributes is essential for augmenting long-tailed datasets and enhancing diagnostic interpretability. However, generic diffusion-based editing methods often induce global distortions or fail to preserve the peri-lesional context critical for diagnosis. To address these limitations, we propose Context-Preserving Dermoscopic Editing (CPDE), a framework tailored for lesion-aware attribute modification. CPDE introduces a dual-branch diffusion pipeline that disentangles lesion editing from background reconstruction. Attribute changes are predicted by a Spatial–channel Transformer at the U-Net bottleneck and are trained with a lesion-aware objective that enforces semantic directionality only inside pathology regions. On the ISIC 2017 and 2018 datasets, CPDE produces spatially localized, clinically coherent edits that preserve lesion extent and surrounding skin. Our method achieves superior image fidelity with an FID of 0.274, while significantly outperforming existing approaches in terms of semantic alignment and background preservation. Yarong Jin, Huanting Guo, Zequn Zhang |
WACV | 5 |
| 2026 | A VAE-GAT-based approach for energy consumption analysis and prediction in manufacturing workshopsabstractIn the global pursuit of carbon neutrality, the manufacturing industry is under increasing pressure to reduce energy waste. Excess consumption not only depletes resources but also hinders sustainable development. Accurate energy consumption prediction is therefore essential for scientific production scheduling and resource allocation, enabling loss reduction, efficiency improvement, and environmental performance enhancement. However, the complexity of modern manufacturing environments results in energy consumption data that is high-dimensional, noisy, and strongly spatiotemporal, which poses challenges to traditional prediction methods. To address these issues, this paper constructs an energy consumption behavior model considering key factors such as equipment status, processing techniques, and environmental conditions. A comprehensive feature analysis and data preprocessing are carried out to identify the key factors influencing consumption. Based on this, an optimization model is proposed that integrates an improved Variational Autoencoder (VAE) with an enhanced Graph Attention Network (GAT). VAE extracts compact latent representations from high-dimensional noisy inputs, suppressing redundancy while preserving essential patterns. GAT then captures complex spatiotemporal dependencies among energy-related features, thereby revealing intrinsic consumption dynamics. Experimental evaluations on both public and real-world datasets demonstrate that the proposed VAE-GAT model achieves superior prediction accuracy and generalization compared with other deep learning baselines. This approach provides a reliable foundation for energy management and contributes to advancing green intelligent manufacturing. Dunbing Tang, Zequn Zhang |
Adv. Eng. Informatics | 5 |
| 2026 | A Large language model-based multi-agent manufacturing system for intelligent shopfloors
Dunbing Tang, Changchun Liu 0002, Liping Wang 0017, Zequn Zhang, Haihua Zhu 0001, Qingwei Nie, Yuchen Ji |
Adv. Eng. Informatics | 5 |
| 2026 | Multi-objective dynamic scheduling in flexible job shops via preference-driven reinforcement learning with meta-path-based transformer model
Jie Chen 0080, Changchun Liu 0002, Zequn Zhang, Dunbing Tang, Liping Wang 0017, Yixiao Jiang |
Expert Syst. Appl. | 4 |
| 2026 | Enhancing dermoscopic image generation via multi-modal conditional diffusion models
Yarong Jin, Huanting Guo, Zequn Zhang, Qian Niu |
Neurocomputing | 5 |
| 2026 | CorrelMamba: Leveraging modality correlations via Bi-Mamba Fusion for brain tumor segmentation
Longgang Yang, Zequn Zhang |
J. Vis. Commun. Image Represent. | 5 |
| 2026 | Research on dynamic obstacle avoidance for industrial AGVs using decay model-based multi-objective Q-learning
Dongdong Li 0006, Dunbing Tang, Zequn Zhang, Lei Wang 0053 |
Knowl. Based Syst. | 4 |
| 2026 | Embodied Intelligence Robots: Flexible Task Planning Framework and Multimodal Fusion Perception
Zequn Zhang, Dunbing Tang, Yuchen Ji |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2026 | Consistency and Invariance Guided Multi-View Hypergraph Learning for Robust Hyperedge PredictionabstractHypergraphs, by extending traditional graphs with hyperedges, enable the modeling and prediction of complex higher-order interactions that go beyond simple pairwise interactions. Hyperedge prediction, an evolution of link prediction, aims to identify potential higher-order interactions—such as those in social media group chats—by recognizing and predicting hyperedges. Recently, hypergraph neural networks (HGNNs) have advanced hyperedge prediction by structuring higher-order interactions into a hypergraph, enabling effective capture of higher-order relations through information propagation across the hypergraph. However, existing methods primarily focus on developing complex HGNNs, underestimating the inherent unreliability of the underlying hypergraph due to incompleteness and noise, leading to suboptimal and fragile performance. In this article, we propose Multi-HyperLinker, a novel multi-view hypergraph learning framework that leverages the consistency and invariance across multiple views to capture reliable higher-order interaction patterns from historical observational data for robust hyperedge prediction. Specifically, to facilitate effective information propagation on incomplete hypergraphs, Multi-HyperLinker first synthesizes a tightly structured hypergraph and designs a consistency-guided dual-view learning strategy. To capture reliable higher-order interaction patterns on noisy hypergraphs, Multi-HyperLinker augments the hypergraphs by perturbing hyperedges to simulate variations and noise, and introduces an invariant learning strategy. Extensive experiments conducted on four real-world datasets demonstrate the superiority of Multi-HyperLinker, achieving performance improvements of up to 19.80% in hit rate compared to existing HGNN-based methods. Additionally, it exhibits enhanced robustness on incomplete and noisy hypergraphs. Changyuan Tian 0001, Li Jin 0001, Zequn Zhang, Zhicong Lu, Wen Shi 0001, Jianhua Yin 0001, Shiyao Yan, Zhi Guo |
ACM Trans. Knowl. Discov. Data | 3 |
| 2025 | HyperMixer: Specializable Hypergraph Channel Mixing for Long-term Multivariate Time Series ForecastingabstractLong-term Multivariate Time Series (LMTS) forecasting aims to predict extended future trends based on channel-interrelated historical data. Considering the elusive channel correlations, most existing methods compromise by treating channels as independent or tentatively modeling pairwise channel interactions, making it challenging to handle the characteristics of both higher-order interactions and time variation in channel correlations. In this paper, we propose HyperMixer, a novel specializable hypergraph channel mixing plugin which introduces versatile hypergraph structures to capture group channel interactions and time-varying patterns for long-term multivariate time series forecasting. Specifically, to encode the higher-order channel interactions, we structure multiple channels into a hypergraph, achieving a two-phase message-passing mechanism: channel-to-group and group-to-channel. Moreover, the functionally specializable hypergraph structures are presented to boost the capability of hypergraph to capture the time-varying patterns across periods, further refining modeling of channel correlations. Extensive experimental results on seven available benchmark datasets demonstrate the effectiveness and generalization of our plugin in LMTS forecasting. The visual analysis further illustrates that HyperMixer with specializable hypergraphs tailors channel interactions specific to certain periods. Changyuan Tian 0001, Zhicong Lu, Zequn Zhang, Heming Yang 0003, Zhi Guo, Xian Sun 0001, Li Jin 0001 |
AAAI | 3 |
| 2025 | RLLTE: Long-Term Evolution Project of Reinforcement LearningabstractWe present RLLTE: a long-term evolution, extremely modular, and open-source framework for reinforcement learning (RL) research and application. Beyond delivering top-notch algorithm implementations, RLLTE also serves as a toolkit for developing algorithms. More specifically, RLLTE decouples the RL algorithms completely from the exploitation-exploration perspective, providing a large number of components to accelerate algorithm development and evolution. In particular, RLLTE is the first RL framework to build a comprehensive ecosystem, which includes model training, evaluation, deployment, benchmark hub, and large language model (LLM)-empowered copilot. RLLTE is expected to set standards for RL engineering practice and be highly stimulative for industry and academia. Our documentation, examples, and source code are available at https://github.com/RLE-Foundation/rllte. Mingqi Yuan, Zequn Zhang, Shihao Luo, Bo Li 0037, Xin Jin 0014, Wenjun Zeng 0001 |
AAAI | 2 |
| 2025 | MLP-Based Vision Language Pre-Training for Medical Visual Question AnsweringabstractMedical visual question answering (VQA) aims to answer clinical questions based on medical images by integrating visual and language information. Due to the limited size of datasets used for medical VQA training, vision-language pretraining (VLP) methods based on the pretrain-finetune paradigm have emerged as the dominant solution in recent years. However, most of the current VLP methods are implemented based on the transformer architecture. The emergence of MLP architecture in the computer vision domain provides us with a novel approach for designing VLP models. In this paper, we propose the first MLP-based vision language pre-training (MBVLP), which is pretrained on three medical image captioning datasets with multiple pre-training objectives and fine-tuned on downstream medical VQA tasks. The proposed MBVLP model surpasses existing methodologies across three medical VQA datasets, achieving state-of-the-art performance. Zequn Zhang |
BIBM | 1 |
| 2025 | Face Relighting with Ratio Function for Explicit Geometric RepresentationabstractThis paper addresses the problem of face relighting under varying illumination conditions. Lighting is a fundamental element in portrait photography that shapes the mood, geometry, and overall realism of the captured characters. Most previous studies have mainly treated relighting as a 2D generation task without incorporating the geometric features of the characters. In contrast, inspired by ratio image-based methods, this paper proposes to disentangle shadow and brightness variations through geometric information and utilizes generative adversarial networks (GANs) to obtain relighted images with brightness consistency. We design a novel relighting-ratio function that integrates the Cook-Torrance reflectance model to more explicitly represent the face geometry than previous ratio image-based methods. This relighting-ratio function is derived from an image rendering formula that quantizes variables such as albedo that are affected by the lighting direction, while systematically excluding variables such as normal and viewpoint that are not affected by lighting. We conduct quantitative and qualitative experiments on the Multi-PIE and CelebA-HQ datasets and show that the proposed method outperforms existing SOTA methods using lighting directions. Yiyang Hu, Zequn Zhang, Hui Zhang 0062, Guquan Jing |
ICASSP | 2 |
| 2025 | HIDE: Hyperspectral Imaging Dataset for Camouflaged Target Recognition
Zequn Zhang, Zhaoyuan Zhang, Zihe Chen, Yangfan Li 0001, Mengquan Li |
ICIC (17) | 4 |
| 2025 | A skill vector-based multi-task optimization algorithm for achieving objectives of multiple users in cloud manufacturing
Yixiao Jiang, Dunbing Tang, Zequn Zhang |
Adv. Eng. Informatics | 6 |
| 2025 | Room-level localization method in industrial workshops using LiDAR-based point cloud registration and object recognition
Libin Tan, Zequn Zhang |
Appl. Intell. | 4 |
| 2025 | How to learn new knowledge: a multimodal contrastive learning framework for open-world knowledge graph completion
Shensi Wang, Kun Fu 0001, Xian Sun 0001, Zequn Zhang, Li Jin 0001, Yuying Shang, Shiyao Yan |
Appl. Intell. | 4 |
| 2025 | HGCMLDA: predicting lncRNA-disease associations using hypergraph contrastive learning and multi-scale attentional feature fusionabstractResearch has consistently indicated that long non-coding RNAs (lncRNAs) significantly influence the development of numerous diseases. Predicting lncRNA-disease associations (LDAs) will contribute to the prevention and treatment of diseases. However, most existing computational models suffer from several challenges: (i) difficulty in capturing complex higher-order relationships among nodes; (ii) limited number of known associations and neglect of consistency of representations across views; (iii) inadequate fusion of multi-view data. In this research, we introduce an innovative end-to-end method named HGCMLDA for LDA prediction. Firstly, HGCMLDA constructs hypergraphs of lncRNAs and diseases based on integrated similarity matrices utilizing Gaussian mixture model and k-nearest neighbor methods and utilizes hypergraph convolutional network to extract high-order representations of lncRNAs and diseases, followed by contrastive learning to capture information interaction between different views that can alleviate the dependence on limited known associations and enhance the node representations in an unsupervised way. Then, HGCMLDA utilizes multi-scale attentional feature fusion, which considers importance weights of different views and aggregates both global and local context to achieve achieve effective and adequate feature fusion. Subsequently, disease features and lncRNA features are also extracted by using variational autoencoder on the association matrix, so that prior knowledge is effectively incorporated for prediction. Finally, the features of these two parts are concatenated, and matrix completion is performed to predict LDA scores. The results of the comparison experiments indicate that HGCMLDA outperforms five state-of-the-art models for LDA prediction. Case studies for specific diseases demonstrate that HGCMLDA can identify novel associations with high accuracy. Zequn Zhang |
Briefings Bioinform. | 1 |
| 2025 | Copiously Quote Classics: Improving Chinese Poetry Generation with historical allusion knowledge
Zhonghe Han, Yuanben Zhang, Zequn Zhang |
Comput. Speech Lang. | 6 |
| 2025 | Self-adaptive production scheduling for discrete manufacturing workshop using multi-agent cyber physical system
Jie Chen 0080, Zequn Zhang, Liping Wang 0017, Dunbing Tang, Qixiang Cai |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | CoSAF: Toward a Secure Meta-Computing IIoT Infrastructure Through Collaborative Source Address FilteringabstractThe rapid growth of the Industrial Internet of Things (IIoT) requires a secure meta-computing environment to support applications like industrial monitoring and remote control. However, this environment faces major security challenges, especially the risk of source address forgery, which can enable DDoS and botnet attacks, disrupting operations and compromising equipment. Current Internet infrastructure forwards packets based only on destination addresses, lacking the capability for source address verification. Although edge-based solutions like firewalls and systems, such as source address validation architecture (SAVA) and source address validation improvement (SAVI), are deployed, they fall short of comprehensive source address validation (SAV), allowing malicious traffic to propagate through core networks. To enhance security, a collaborative approach based on meta-computing principles is needed, allowing routers to verify source addresses cooperatively. Given the impracticality of fully upgrading routers, incremental deployment is essential. We show that optimizing incremental SAV deployment is NP-hard. To address this, we propose collaborative optimized source address filtering (COSAF), a heuristic algorithm that uses a sink-tree structure to effectively filter attack flows and optimize resource allocation. COSAF also takes SAV table capacity into account to improve resource utilization. Extensive simulations demonstrate that COSAF outperforms traditional methods. Shu Yang 0002, Zequn Zhang, Laizhong Cui |
IEEE Internet Things J. | 2 |
| 2025 | Dynamic scheduling for dual-resource constrained flexible job-shop via semantic-aware graph modelling and deep reinforcement learning
Jie Chen 0080, Zequn Zhang, Dunbing Tang, Qixiang Cai |
Knowl. Based Syst. | 3 |
| 2025 | RAIN: Reconstructed-aware in-context enhancement with graph denoising for session-based recommendation
Xinyi Zeng, Shuchao Li, Zequn Zhang, Li Jin 0001, Zhi Guo, Kaiwen Wei |
Neural Networks | 3 |
| 2025 | Canvas: Compositional Generation for Art Painting With Seamless Subject-Driven Infusion
Yunnan Wang, Lexiang Lv, Zequn Zhang, Xiaoyu Shen 0001, Xin Jin 0014, Wenjun Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Exploring Contrastive Pre-Training for Domain Connections in Medical Image SegmentationabstractUnsupervised domain adaptation (UDA) in medical image segmentation aims to improve the generalization of deep models by alleviating domain gaps caused by inconsistency across equipment, imaging protocols, and patient conditions. However, existing UDA works remain insufficiently explored and present great limitations: 1) Exhibit cumbersome designs that prioritize aligning statistical metrics and distributions, which limits the model's flexibility and generalization while also overlooking the potential knowledge embedded in unlabeled data; 2) More applicable in a certain domain, lack the generalization capability to handle diverse shifts encountered in clinical scenarios. To overcome these limitations, we introduce MedCon, a unified framework that leverages general unsupervised contrastive pre-training to establish domain connections, effectively handling diverse domain shifts without tailored adjustments. Specifically, it initially explores a general contrastive pre-training to establish domain connections by leveraging the rich prior knowledge from unlabeled images. Thereafter, the pre-trained backbone is fine-tuned using source-based images to ultimately identify per-pixel semantic categories. To capture both intra- and inter-domain connections of anatomical structures, we construct positive-negative pairs from a hybrid aspect of both local and global scales. In this regard, a shared-weight encoder-decoder is employed to generate pixel-level representations, which are then mapped into hyper-spherical space using a non-learnable projection head to facilitate positive pair matching. Comprehensive experiments on diverse medical image datasets confirm that MedCon outperforms previous methods by effectively managing a wide range of domain shifts and showcasing superior generalization capabilities. Zequn Zhang, Yunnan Wang, Baao Xie, Yuhang Li 0005, Zhen Chen 0013, Xin Jin 0014, Wenjun Zeng 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Flexible Optimal Transport With Contrastive Graphical Modeling for Multimodal Hate DetectionabstractMultimodal hate detection plays a crucial role in maintaining harmonious online environments by identifying harmful content, such as hateful memes. Although previous research has made significant progress in detecting explicit hate speech, there remains a critical gap in analyzing implicit hate, which is particularly challenging due to the absence of explicit harmful text claims or demographic visual cues. Despite the promising results based on cross-modal attention, previous methods may suffer from the distributional modality gap caused by the non-literal associations between multimodal elements, which lacks apparent alignment in implicit hateful contents. In this work, we propose a novel framework: Flexible Optimal Transport (FLOT) to capture the non-literal cross-modal alignment for multimodal hate in the context of memes. FLOT formulates the problem of cross-modal alignment as finding optimal transportation plans, which leverages a kernel method to capture complementary information from multiple modalities. The kernel embeddings reproduce a kernel Hilbert space (RKHS) to serve as a non-linear transformation of alignment, which effectively reduces the distributional modality gap with more interpretability. Moreover, we established topological structures with contrastive modeling for the aligned representations, which are optimized to achieve comprehensive alignment between different modalities, and facilitate local reasoning based on multimodal elements. Experimental results have demonstrated that our FLOT achieved state-of-the-art performance on three publicly available benchmark datasets. Furthermore, extensive qualitative analysis confirms the superior ability of FLOT in capturing implicit cross-modal alignment. Linhao Zhang, Li Jin 0001, Xiaoyu Li 0004, Xian Sun 0001, Xin Wang 0117, Zequn Zhang, Jian Liu 0032, Zhicong Lu, Guangluan Xu |
IEEE Trans. Multim. | 6 |
| 2025 | Unleash the Power of Vision-Language Models by Visual Attention Prompt and Multimodal InteractionabstractPre-trained vision-language models (VLMs), equipped with parameter-efficient tuning (PET) methods like prompting, have shown impressive knowledge transferability on new downstream tasks, but they are still prone to be limited by catastrophic forgetting and overfitting dilemma due to large gaps among tasks. Furthermore, the underlying physical mechanisms of prompt-based tuning methods (especially for visual prompting) remain largely unexplored. It is unclear why these methods work solely based on learnable parameters as prompts for adaptation. To address the above challenges, we present a new prompt-based framework for vision-language models, termed Uni-prompt. Our framework transfers VLMs to downstream tasks by designing visual prompts from an attention perspective that reduces the transfer/solution space, which enables the vision model to focus on task-relevant regions of the input image while also learning task-specific knowledge. Additionally, Uni-prompt further aligns visual-text prompts learning through a pretext task with masked representation modeling interactions, which implicitly learns a global cross-modal matching between visual and language concepts for consistency. We conduct extensive experiments on the few-shot classification task and achieve significant improvement using our Uni-prompt method while requiring minimal extra parameters cost. Letian Wu, Zequn Zhang, Tao Yu 0012, Chao Ma 0004, Xin Jin 0014, Xiaokang Yang 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | scMonica: Single-cell Mosaic Omics Nonlinear Integration and Clustering AnalysisabstractSingle-cell mosaic integration has revolutionized our understanding of cellular heterogeneity, offering unprecedented resolution into cellular states and contexts. While large language models (LLMs) have achieved success in the analysis of single-cell omics data, their application specifically focused on the integration of mosaic data remains limited. Current computational approaches often fall short in addressing these challenges. They frequently depend on assumptions of data completeness and uniform quality, failing to manage the variability and noise effectively introduced by missing data. Moreover, these models might not efficiently process large, diverse datasets or account for biological heterogeneity and long-ranging dependencies across various data types. To address these shortcomings, we introduce the Single-cell Mosaic Omics Nonlinear Integration and Clustering Analysis (scMonica) framework, which employs a LSTM-transformer hybrid architecture. This innovative model combines the strengths of Long Short-Term Memory (LSTM) networks, which excel at capturing long-range dependencies within sequential gene expression patterns, with transformers, renowned for their attention mechanisms that handle the complex, non-linear interactions characteristic of multi-layered datasets. By leveraging these complementary strengths, our approach enhances the integration process significantly, allowing for nuanced management of the intrinsic heterogeneity and sparsity of mosaic datasets. Comprehensive evaluations demonstrate the robustness and effectiveness of our approach, offering unparalleled versatility and accuracy in multi-omics data analysis. These advancements underscore scMonica’s potential to drive significant insights in single-cell developmental biology, oncology, and beyond. We discuss the underlying technologies, analyze their applications, and contemplate future directions that promise to extend the boundaries of both research and clinical domains. Saba Aslam, Zequn Zhang, Ruey-Song Huang |
BIBM | 6 |
| 2024 | Consistency Prior Matters: Biomedical-Prompting Dual Augmentation for Domain Adaptive Medical Image SegmentationabstractExisting domain adaptive medical image segmentation works typically rely on style transfer techniques to mitigate the unexpected domain gap, which inevitably suffers from synthesized artifacts or unreasonable stylization. In this paper, we propose to inject biomedical-related prior knowledge (i.e., intensity and anatomical consistency) as regularization in a prompting manner, bridging the domain gap across modalities. Technically, we develop an efficient scheme called Biomedical-Prompting Dual Augmentation (BPDA) to learn domain-invariant representations by enforcing consistent model predictions across different augmented views. BPDA augments unpaired source and target images from intensity and anatomical aspects in a dual manner, while prompting the framework to fully understand the anatomical structure-invariant features. In this way, our method captures discriminative inherent representations on cross-modality scenarios. Furthermore, we also introduce a Cross-Domain Prototype Denoising (CDPD) in BPDA to refine pseudo-labeling results with the class centroids for a reliable augmentation. Extensive experiments on the cross-modality abdominal and cardiac segmentation benchmarks demonstrate the superiority of our method over state-of-the-art alternatives. Yunnan Wang, Zequn Zhang, Xin Jin 0014, Wenjun Zeng 0001 |
BIBM | 2 |
| 2024 | Spiking Generative Adversarial Network for Controllable Affective Music Creation
Xianghong Lin, Zequn Zhang, Chengyang Xie, Ruidong Ma |
ICIC (2) | 3 |
| 2024 | Scene Graph Disentanglement and Composition for Generalizable Complex Image GenerationabstractThere has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their relationships. To address this issue, we leverage the scene graph, a powerful structured representation, for complex image generation. Different from the previous works that directly use scene graphs for generation, we employ the generative capabilities of variational autoencoders and diffusion models in a generalizable manner, compositing diverse disentangled visual clues from scene graphs. Specifically, we first propose a Semantics-Layout Variational AutoEncoder (SL-VAE) to jointly derive (layouts, semantics) from the input scene graph, which allows a more diverse and reasonable generation in a one-to-many mapping. We then develop a Compositional Masked Attention (CMA) integrated with a diffusion model, incorporating (layouts, semantics) with fine-grained attributes as generation guidance. To further achieve graph manipulation while keeping the visual content consistent, we introduce a Multi-Layered Sampler (MLS) for an "isolated" image editing effect. Extensive experiments demonstrate that our method outperforms recent competitors based on text, layout, or scene graph, in terms of generation rationality and controllability. Yunnan Wang, Zequn Zhang, Baao Xie, Xihui Liu, Wenjun Zeng 0001, Xin Jin 0014 |
NeurIPS | 4 |
| 2024 | Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language ModelsabstractDisentangled representation learning (DRL) aims to identify and decompose underlying factors behind observations, thus facilitating data perception and generation. However, current DRL approaches often rely on the unrealistic assumption that semantic factors are statistically independent. In reality, these factors may exhibit correlations, which off-the-shelf solutions have yet to properly address. To tackle this challenge, we introduce a bidirectional weighted graph-based framework, to learn factorized attributes and their interrelations within complex data. Specifically, we propose a $\beta$-VAE based module to extract factors as the initial nodes of the graph, and leverage the multimodal large language model (MLLM) to discover and rank latent correlations, thereby updating the weighted edges. By integrating these complementary modules, our model successfully achieves fine-grained, practical and unsupervised disentanglement. Experiments demonstrate our method's superior performance in disentanglement and reconstruction. Furthermore, the model inherits enhanced interpretability and generalizability from MLLMs. Baao Xie, Qiuyu Chen, Yunnan Wang, Zequn Zhang, Xin Jin 0014, Wenjun Zeng 0001 |
NeurIPS | 4 |
| 2024 | Collaborative dynamic scheduling in a self-organizing manufacturing system using multi-agent reinforcement learning
Yong Gui, Zequn Zhang, Dunbing Tang, Haihua Zhu 0001, Yi Zhang 0136 |
Adv. Eng. Informatics | 2 |
| 2024 | Probing a point cloud based expeditious approach with deep learning for constructing digital twin models in shopfloor
Zequn Zhang, Qingwei Nie, Dunbing Tang |
Adv. Eng. Informatics | 2 |
| 2024 | Graph-enhanced context aware framework for session-based recommendation
Xinyi Zeng, Zequn Zhang, Shuchao Li, Zhi Guo, Li Jin 0001, Xian Sun 0001 |
Neurocomputing | 2 |
| 2024 | SKYPER: Legal case retrieval via skeleton-aware hypergraph embedding in the hyperbolic space
Shiyao Yan, Zequn Zhang |
Inf. Sci. | 2 |
| 2024 | Dense-sparse representation matters: A point-based method for volumetric medical image segmentation
Bingxi Liu 0003, Zequn Zhang, Yao Yan 0003, Huanting Guo |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | LollipopE: Bi-centered lollipop embedding for complex logic query on knowledge graph
Shiyao Yan, Changyuan Tian 0001, Zequn Zhang, Guangluan Xu |
Neural Networks | 3 |
| 2024 | More Than Syntaxes: Investigating Semantics to Zero-shot Cross-lingual Relation Extraction and Event Argument Role LabellingabstractSyntactic dependency structures are commonly utilized as language-agnostic features to solve the word order difference issues in zero-shot cross-lingual relation and event extraction tasks. However, while sentences in multiple forms can be employed to express the same meaning, the syntactic structure may vary considerably in specific scenarios. To fix this problem, we find semantics are rarely considered, which could provide a more consistent semantic analysis of sentences and be served as another bridge between different languages. Therefore, in this article, we introduce Syntax and Semantic Driven Network (SSDN) to equip syntax and semantic knowledge across languages simultaneously. Specifically, predicate–argument structures from semantic role labelling are explicitly incorporated into word representations. Then, a semantic-aware relational graph convolutional network and a transformer-based encoder are utilized to model both semantic dependency and syntactic dependency structures, respectively. Finally, a fusion module is introduced to integrate output representations adaptively. We conduct experiments on the widely used Automatic Content Extraction 2005 English, Chinese, and Arabic datasets. The evaluation results demonstrate that the proposed method achieves the state-of-the-art performance. Further study also indicates SSDN could produce robust representations that facilitate the transfer operations across languages. Kaiwen Wei, Li Jin 0001, Zequn Zhang, Zhi Guo, Xiaoyu Li 0004, Qing Liu 0021, Weimiao Feng |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | TOT:Topology-Aware Optimal Transport for Multimodal Hate DetectionabstractMultimodal hate detection, which aims to identify the harmful content online such as memes, is crucial for building a wholesome internet environment. Previous work has made enlightening exploration in detecting explicit hate remarks. However, most of their approaches neglect the analysis of implicit harm, which is particularly challenging as explicit text markers and demographic visual cues are often twisted or missing. The leveraged cross-modal attention mechanisms also suffer from the distributional modality gap and lack logical interpretability. To address these semantic gap issues, we propose TOT: a topology-aware optimal transport framework to decipher the implicit harm in memes scenario, which formulates the cross-modal aligning problem as solutions for optimal transportation plans. Specifically, we leverage an optimal transport kernel method to capture complementary information from multiple modalities. The kernel embedding provides a non-linear transformation ability to reproduce a kernel Hilbert space (RKHS), which reflects significance for eliminating the distributional modality gap. Moreover, we perceive the topology information based on aligned representations to conduct bipartite graph path reasoning. The newly achieved state-of-the-art performance on two publicly available benchmark datasets, together with further visual analysis, demonstrate the superiority of TOT in capturing implicit cross-modal alignment. Linhao Zhang, Li Jin 0001, Xian Sun 0001, Guangluan Xu, Zequn Zhang, Xiaoyu Li 0004, Nayu Liu, Qing Liu 0021, Shiyao Yan |
AAAI | 5 |
| 2023 | Guide the Many-to-One Assignment: Open Information Extraction via IoU-aware Optimal TransportabstractKaiwen Wei, Yiran Yang, Li Jin, Xian Sun, Zequn Zhang, Jingyuan Zhang, Xiao Li, Linhao Zhang, Jintao Liu, Guo Zhi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kaiwen Wei, Li Jin 0001, Xian Sun 0001, Zequn Zhang, Linhao Zhang, Zhi Guo |
ACL (1) | 5 |
| 2023 | SRNet: Striped Pyramid Pooling and Relational Transformer for Retinal Vessel Segmentation
Zequn Zhang, Yao Yan 0003, Bingxi Liu 0003 |
BMVC | 3 |
| 2023 | Event Causality Extraction via Implicit Cause-Effect InteractionsabstractEvent Causality Extraction (ECE) aims to extract the cause-effect event pairs from the given text, which requires the model to possess a strong reasoning ability to capture event causalities.However, existing works have not adequately exploited the interactions between the cause and effect event that could provide crucial clues for causality reasoning.To this end, we propose an Implicit Cause-Effect interaction (ICE) framework, which formulates ECE as a template-based conditional generation problem.The proposed method captures the implicit intra-and inter-event interactions by incorporating the privileged information (ground truth event types and arguments) for reasoning, and a knowledge distillation mechanism is introduced to alleviate the unavailability of privileged information in the test stage.Furthermore, to facilitate knowledge transfer from teacher to student, we design an event-level alignment strategy named Cause-Effect Optimal Transport (CEOT) to strengthen the semantic interactions of cause-effect event types and arguments.Experimental results indicate that ICE achieves state-of-the-art performance on the ECE-CCKS dataset. Zequn Zhang, Kaiwen Wei, Zhi Guo, Xian Sun 0001, Li Jin 0001, Xiaoyu Li 0004 |
EMNLP | 2 |
| 2023 | NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigationabstract3D representation disentanglement aims to identify, decompose, and manipulate the underlying explanatory factors of 3D data, which helps AI fundamentally understand our 3D world. This task is currently under-explored and poses great challenges: (i) the 3D representations are complex and in general contains much more information than 2D image; (ii) many 3D representations are not well suited for gradient-based optimization, let alone disentanglement. To address these challenges, we use NeRF as a differentiable 3D representation, and introduce a self-supervised Navigation to identify interpretable semantic directions in the latent space. To our best knowledge, this novel method, dubbed NaviNeRF, is the first work to achieve fine-grained 3D disentanglement without any priors or supervisions. Specifically, NaviNeRF is built upon the generative NeRF pipeline, and equipped with an Outer Navigation Branch and an Inner Refinement Branch. They are complementary —— the outer navigation is to identify global-view semantic directions, and the inner refinement dedicates to fine-grained attributes. A synergistic loss is further devised to coordinate two branches. Extensive experiments demonstrate that NaviNeRF has a superior fine-grained 3D disentanglement ability than the previous 3D-aware models. Its performance is also comparable to editing-oriented models relying on semantic or geometry priors.* Baao Xie, Bohan Li 0015, Zequn Zhang, Junting Dong, Xin Jin 0014, Jing-Yu Yang 0002, Wenjun Zeng 0001 |
ICCV | 3 |
| 2023 | Emotion-cause pair extraction with bidirectional multi-label sequence tagging
Zequn Zhang, Zhi Guo, Li Jin 0001, Xiaoyu Li 0004, Kaiwen Wei, Xian Sun 0001 |
Appl. Intell. | 2 |
| 2023 | MDSC-Net: A multi-scale depthwise separable convolutional neural network for skin lesion segmentationabstractAbstract Accurate segmentation of the skin lesion region is crucial for diagnosing and screening skin diseases. However, skin lesion segmentation is challenging due to the indistinguishable boundaries of the lesion region, irregular shapes and hair interference. To settle the above issues, we propose a Multi‐scale Depthwise Separable Convolutional Neural Network for skin lesion segmentation named MDSC‐Net. Specifically, a novel Multi‐scale Depthwise Separable Residual Convolution Module is employed in skip connection, conveying more detailed features to the decoder. To compensate for the loss of spatial location information in down‐sampling, we propose a novel Spatial Adaption Module. Furthermore, we propose a Multi‐scale Decoding Fusion Module in the decoder to capture contextual information. We have performed extensive experiments to verify the effectiveness and robustness of the proposed network on three public benchmark skin lesion segmentation datasets and one public benchmark polyp segmentation dataset, including ISIC‐2017, ISIC‐2018, PH2, and Kvasir‐SEG datasets. Experimental results consistently demonstrate the proposed MDSC‐Net achieves superior segmentation across five popularly used evaluation criteria. The proposed network reaches high‐performance skin lesion segmentation, and can provide important clues to help doctors diagnose and treat skin cancer early. Zequn Zhang |
IET Image Process. | 3 |
| 2023 | ReasonFuse: Reason Path Driven and Global-Local Fusion Network for Numerical Table-Text Question Answering
Yuancheng Xia, Feng Li 0030, Qing Liu 0021, Li Jin 0001, Zequn Zhang, Xian Sun 0001, Lixu Shao |
Neurocomputing | 5 |
| 2023 | KEPT: Knowledge Enhanced Prompt Tuning for event causality identification
Zequn Zhang, Zhi Guo, Li Jin 0001, Xiaoyu Li 0004, Kaiwen Wei, Xian Sun 0001 |
Knowl. Based Syst. | 2 |
| 2023 | Exploiting event-aware and role-aware with tree pruning for document-level event extraction
Jianwei Lv, Zequn Zhang, Guangluan Xu, Xian Sun 0001, Shuchao Li, Qing Liu 0021, Pengcheng Dong |
Neural Comput. Appl. | 2 |
| 2023 | Tackling higher-order relations and heterogeneity: Dynamic heterogeneous hypergraph network for spatiotemporal activity prediction
Changyuan Tian 0001, Zequn Zhang, Fanglong Yao, Zhi Guo, Shiyao Yan, Xian Sun 0001 |
Neural Networks | 2 |
| 2023 | Extracting 3-D Structural Lines of Building From ALS Point Clouds Using Graph Neural Network Embedded With Corner InformationabstractThe representation quantifies the geometric shape and topology of a building is a necessary procedure for many urban planning applications. A sharp line framework is a high-level structural cue providing a compact building representation. However, accurate and efficient structural line extraction remains a challenging task given the variety and complexity of buildings. This study proposes a general 3-D structural line extraction method from point clouds. The building points are extracted and further divided into various single-building units. In the proposed 3-D structural line extraction method, individual building point cloud is the input. First, the corners are detected by an associative learning module. Next, the curve connection is implemented by a link prediction block based on the graph neural network (GNN) embedded with corner information. After that, the obtained curves are subsequently converted into a topological graph. Finally, the corner points are optimized to achieve precise fitting of the structural lines. The experiments and comparisons on two airborne laser scanning (ALS) point cloud datasets demonstrate the effectiveness of the proposed method and the ability to retrieve ideal structural line results for building point clouds. Furthermore, without reprocessing, the proposed method yielded better results for various dataset types (outdoor building, indoor scene, and furniture point clouds) than the prevalent published methods (i.e., EC-Net, PIE-Net, and PC2WF), verifying its strength and efficacy. To further verify the accuracy of the obtained structural lines, we also introduce a line-based model reconstruction method that employ these lines for building reconstruction. Tengping Jiang, Zequn Zhang, Yongchao Yang, Xin Jin 0014, Wenjun Zeng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Implicit Event Argument Extraction With Argument-Argument Relational KnowledgeabstractAs a challenging sub-task of event argument extraction, implicit event argument extraction seeks to identify document-level arguments that play direct or implicit roles in a given event. Prior work mainly focuses on capturing direct relations between arguments and the event trigger; however, the lack of reasoning ability imposes limitations to the extraction of implicit arguments. In this work, we propose anArgument-argumentRelation-enhancedEventArgument extraction (AREA) learning framework to tackle this issue through reasoning in event frame-level scope. The proposed method leverages related arguments of the expected one as clues, and utilizes such argument-argument dependencies to guide the reasoning process. To bridge the distribution gap between oracle knowledge used in the training phase and the imperfect related arguments in the test stage, we introduce a conventional knowledge distillation strategy to drive a final model that can work without extra inputs by mimicking the behaviour of a well-informed teacher model. In addition, considering that conventional knowledge distillation methods transfer knowledge individually, we integrate it with a novel relational knowledge distillation mechanism to explicitly capture the structural mutual argument-argument relation. Moreover, since the training process is not compatible with the real situation, a curriculum learning method is further introduced to make the training process smoother. Experimental results demonstrate that the learning framework obtains state-of-the-art performance on the RAMS and Wikievents datasets. Ablation study and further discussion also show it could handle long-range dependency and implicit argument problems effectively. Kaiwen Wei, Xian Sun 0001, Zequn Zhang, Li Jin 0001, Jianwei Lv, Zhi Guo |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | PolygonE: Modeling N-ary Relational Data as Gyro-Polygons in Hyperbolic SpaceabstractN-ary relational knowledge base (KBs) embedding aims to map binary and beyond-binary facts into low-dimensional vector space simultaneously. Existing approaches typically decompose n-ary relational facts into subtuples (entity pairs, triples or quintuples, etc.), and they generally model n-ary relational KBs in Euclidean space. However, n-ary relational facts are semantically and structurally intact, decomposition leads to the loss of global information and undermines the semantical and structural integrity. Moreover, compared to the binary relational KBs, n-ary ones are characterized by more abundant and complicated hierarchy structures, which could not be well expressed in Euclidean space. To address the issues, we propose a gyro-polygon embedding approach to realize n-ary fact integrity keeping and hierarchy capturing, termed as PolygonE. Specifically, n-ary relational facts are modeled as gyro-polygons in the hyperbolic space, where we denote entities in facts as vertexes of gyro-polygons and relations as entity translocation operations. Importantly, we design a fact plausibility measuring strategy based on the vertex-gyrocentroid geodesic to optimize the relation-adjusted gyro-polygon. Extensive experiments demonstrate that PolygonE shows SOTA performance on all benchmark datasets, generalizability to binary data, and applicability to arbitrary arity fact. Finally, we also visualize the embedding to help comprehend PolygonE's awareness of hierarchies. Shiyao Yan, Zequn Zhang, Xian Sun 0001, Guangluan Xu, Shuchao Li, Qing Liu 0021, Nayu Liu, Shensi Wang |
AAAI | 2 |
| 2022 | Unsupervised Heterogeneous Cryo-EM Projection Image Classification Using AutoencoderabstractHeterogeneous three-dimensional (3D) reconstruction in single-particle cryo-electron microscopy (cryo-EM) is a significant but very challenging technique for recovering conformational heterogeneity of proteins or other biological macromolecules and their complexes in different functional states. Heterogeneous projection image classification is an effective way for solving the heterogeneity problem in single-particle cryo-EM. Most existing heterogeneous projection image classification methods are based on supervised learning or require a large amount of a priori knowledge, such as the common lines or orientations of the projection images, which has many limitations in practical applications. In this paper, we propose an unsupervised heterogeneous cryo-EM projection image classification algorithm based on autoencoders that only needs to know the number of heterogeneous conformations in the dataset and does not require any labeling information of the projection images as well as other prior knowledge. We implement a simple autoencoder with a multi-layer perceptron that is trained in iterative mode and a complex autoencoder with a residual network that is trained in one-pass learning mode to convert heterogeneous projection images into latent variables. The extracted high-dimensional features are reduced to two dimensions by the uniform manifold approximate and projection dimensionality reduction algorithm and then cluster them using the spectral clustering algorithm. The proposed algorithm is applied to two heterogeneous cryo-EM datasets to demonstrate its classification performance. Experimental results show that the proposed algorithm can effectively extract category features of heterogeneous projection images and can classify them with high accuracy. Yonggang Lu, Zequn Zhang |
BIBM | 4 |
| 2022 | A Span-level Bidirectional Network for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE)is a new fine-grained sentiment analysis task that aims to extract triplets of aspect terms, sentiments, and opinion terms from review sentences.Recently, span-level models achieve gratifying results on ASTE task by taking advantage of the predictions of all possible spans.Since all possible spans significantly increases the number of potential aspect and opinion candidates, it is crucial and challenging to efficiently extract the triplet elements among them.In this paper, we present a span-level bidirectional network which utilizes all possible spans as input and extracts triplets from spans bidirectionally.Specifically, we devise both the aspect decoder and opinion decoder to decode the span representations and extract triples from aspect-to-opinion and opinion-to-aspect directions.With these two decoders complementing with each other, the whole network can extract triplets from spans more comprehensively.Moreover, considering that mutual exclusion cannot be guaranteed between the spans, we design a similar span separation loss to facilitate the downstream task of distinguishing the correct span by expanding the KL divergence of similar spans during the training process; in the inference process, we adopt an inference strategy to remove conflicting triplets from the results base on their confidence scores.Experimental results show that our framework not only significantly outperforms state-of-the-art methods, but achieves better performance in predicting triplets with multi-token entities and extracting triplets in sentences contain multitriplets 1 . Yuqi Chen 0015, Zequn Zhang |
EMNLP | 4 |
| 2022 | DPNet: domain-aware prototypical network for interdisciplinary few-shot relation classification
Li Jin 0001, Xiaoyu Li 0004, Xian Sun 0001, Zhi Guo, Zequn Zhang, Shuchao Li |
Appl. Intell. | 6 |
| 2022 | SF-ANN: leveraging structural features with an attention neural network for candidate fact ranking
Li Jin 0001, Zequn Zhang, Xiaoyu Li 0004, Qing Liu 0021 |
Appl. Intell. | 3 |
| 2022 | RSAP-Net: joint optic disc and cup segmentation with a residual spatial attention path module and MSRCR-PT pre-processing algorithmabstractBACKGROUND: Glaucoma can cause irreversible blindness to people's eyesight. Since there are no symptoms in its early stage, it is particularly important to accurately segment the optic disc (OD) and optic cup (OC) from fundus medical images for the screening and prevention of glaucoma. In recent years, the mainstream method of OD and OC segmentation is convolution neural network (CNN). However, most existing CNN methods segment OD and OC separately and ignore the a priori information that OC is always contained inside the OD region, which makes the segmentation accuracy of most methods not high enough. METHODS: This paper proposes a new encoder-decoder segmentation structure, called RSAP-Net, for joint segmentation of OD and OC. We first designed an efficient U-shaped segmentation network as the backbone. Considering the spatial overlap relationship between OD and OC, a new Residual spatial attention path is proposed to connect the encoder-decoder to retain more characteristic information. In order to further improve the segmentation performance, a pre-processing method called MSRCR-PT (Multi-Scale Retinex Colour Recovery and Polar Transformation) has been devised. It incorporates a multi-scale Retinex colour recovery algorithm and a polar coordinate transformation, which can help RSAP-Net to produce more refined boundaries of the optic disc and the optic cup. RESULTS: The experimental results show that our method achieves excellent segmentation performance on the Drishti-GS1 standard dataset. In the OD and OC segmentation effects, the F1 scores are 0.9752 and 0.9012, respectively. The BLE are 6.33 pixels and 11.97 pixels, respectively. CONCLUSIONS: This paper presents a new framework for the joint segmentation of optic discs and optic cups, called RSAP-Net. The framework mainly consists of a U-shaped segmentation skeleton and a residual space attention path module. The design of a pre-processing method called MSRCR-PT for the OD/OC segmentation task can improve segmentation performance. The method was evaluated on the publicly available Drishti-GS1 standard dataset and proved to be effective. Zeqi Ma, Zequn Zhang |
BMC Bioinform. | 4 |
| 2022 | Span-based dual-decoder framework for aspect sentiment triplet extraction
Yuqi Chen 0015, Zequn Zhang |
Neurocomputing | 2 |
| 2022 | TSPNet: Translation supervised prototype network via residual learning for multimodal social relation extraction
Hankun Kang, Xiaoyu Li 0004, Li Jin 0001, Zequn Zhang, Shuchao Li |
Neurocomputing | 5 |
| 2022 | Representation learning of knowledge graphs with the interaction between entity types and relations
Shensi Wang, Kun Fu 0001, Xian Sun 0001, Zequn Zhang, Shuchao Li, Shiyao Yan |
Neurocomputing | 4 |
| 2022 | HYPER2: Hyperbolic embedding for hyper-relational link prediction
Shiyao Yan, Zequn Zhang, Xian Sun 0001, Guangluan Xu, Li Jin 0001, Shuchao Li |
Neurocomputing | 2 |
| 2022 | Trigger is Non-central: Jointly event extraction via label-aware representations with multi-task learning
Jianwei Lv, Zequn Zhang, Li Jin 0001, Shuchao Li, Xiaoyu Li 0004, Guangluan Xu, Xian Sun 0001 |
Knowl. Based Syst. | 2 |
| 2022 | HEFT: A History-Enhanced Feature Transfer framework for incremental event detection
Kaiwen Wei, Zequn Zhang, Li Jin 0001, Zhi Guo, Shuchao Li, Jianwei Lv |
Knowl. Based Syst. | 2 |
| 2022 | Modeling N-ary relational data as gyro-polygons with learnable gyro-centroid
Shiyao Yan, Zequn Zhang, Guangluan Xu, Xian Sun 0001, Shuchao Li, Shensi Wang |
Knowl. Based Syst. | 2 |
| 2021 | Trigger is Not Sufficient: Exploiting Frame-aware Knowledge for Implicit Event Argument ExtractionabstractKaiwen Wei, Xian Sun, Zequn Zhang, Jingyuan Zhang, Guo Zhi, Li Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kaiwen Wei, Xian Sun 0001, Zequn Zhang, Zhi Guo, Li Jin 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | HGEED: Hierarchical graph enhanced event detection
Jianwei Lv, Zequn Zhang, Li Jin 0001, Shuchao Li, Xiaoyu Li 0004, Guangluan Xu, Xian Sun 0001 |
Neurocomputing | 2 |
| 2021 | Hierarchical-aware relation rotational knowledge graph embedding for link prediction
Shensi Wang, Kun Fu 0001, Xian Sun 0001, Zequn Zhang, Shuchao Li, Li Jin 0001 |
Neurocomputing | 4 |
| 2021 | A unified position-aware convolutional neural network for aspect based sentiment analysis
Feng Li 0030, Zequn Zhang, Guangluan Xu, Xian Sun 0001 |
Neurocomputing | 3 |
| 2021 | End-to-end aspect-based sentiment analysis with hierarchical multi-task learning
Guangluan Xu, Zequn Zhang, Li Jin 0001, Xian Sun 0001 |
Neurocomputing | 3 |
| 2021 | Integrate syntax information for target-oriented opinion words extraction with target-specific graph convolutional network
Feng Li 0030, Zequn Zhang, Guangluan Xu, Yang Wang 0056, Yunyan Zhang |
Neurocomputing | 3 |
| 2018 | High Resolution SAR Image Classification with Deeper Convolutional Neural NetworkabstractDeeper architectures are proven to be beneficial for the classification performance obviously in computer vision field. Inspired by this, deep CNN s are expected to make progress in the SAR target classification problem as well. However, it is hard to train deeper CNNs for SAR images. Such CNNs have millions of parameters to be determined in the network (for example the VGGNet has more than 130 million parameters), hence large-scale dataset is indispensable when training a deep CNN. But there is no large-scale annotated SAR target dataset, and data acquisition and annotation is much more costly for SAR images. With inadequate data, the network is easy to be overfitting. Several methods based on deep learning have been proposed for SAR image classifications, but they cannot get rid of the aforementioned data limitation of labelled SAR images. To solve this problem, this paper proposes a microarchitecture called CompressUnit (CU). With CU, we design a deeper CNN. Compared with the network with the fewest parameters for SAR image classification in literature so far, our network is 2X deeper with only about 10% of parameters. In this way, we get a deeper network with much fewer parameters. This network is easier to be trained with limited SAR data and is more likely to get rid of overfitting. Yue Zhang 0016, Xian Sun 0001, Hao Sun 0009, Zequn Zhang, Wenhui Diao, Kun Fu 0001 |
IGARSS | 4 |
| 2013 | A direct mining approach to efficient constrained graph pattern discoveryabstractDespite the wealth of research on frequent graph pattern mining, how to efficiently mine the complete set of those with constraints still poses a huge challenge to the existing algorithms mainly due to the inherent bottleneck in the mining paradigm. In essence, mining requests with explicitly-specified constraints cannot be handled in a way that is direct and precise. In this paper, we propose a direct mining framework to solve the problem and illustrate our ideas in the context of a particular type of constrained frequent patterns --- the "skinny" patterns, which are graph patterns with a long backbone from which short twigs branch out. These patterns, which we formally define as l-long δ-skinny patterns, are able to reveal insightful spatial and temporal trajectory patterns in mobile data mining, information diffusion, adoption propagation, and many others. Feida Zhu 0001, Zequn Zhang, Qiang Qu 0001 |
SIGMOD Conference | 2 |
| 2012 | Classification-Based Prediction on the Retweet Actions over Microblog Dataset
Lianshuai Zhang, Zequn Zhang, Peiquan Jin |
WISE | 2 |