Meina Song

dblp:95/4440 · DBLP profile ↗
← Back
53ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0001-6626-9932ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Systems, architecture and hardware · 5Databases, data management, data science and information retrieval · 5 · 4 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 Structures Meet Semantics: Multimodal Fusion via Graph Contrastive Learning
abstract
Multimodal sentiment analysis (MSA) aims to infer emotional states by effectively integrating textual, acoustic, and visual modalities. Despite notable progress, existing multimodal fusion methods often neglect modality-specific structural dependencies and semantic misalignment, limiting their quality, interpretability, and robustness. To address these challenges, we propose a novel framework called the Structural-Semantic Unifier (SSU), which systematically integrates modality-specific structural information and cross-modal semantic grounding for enhanced multimodal representations. Specifically, SSU dynamically constructs modality-specific graphs by leveraging linguistic syntax for text and a lightweight, text-guided attention mechanism for acoustic and visual modalities, thus capturing detailed intra-modal relationships and semantic interactions. We further introduce a semantic anchor, derived from global textual semantics, that serves as a cross-modal alignment hub, effectively harmonizing heterogeneous semantic spaces across modalities. Additionally, we develop a multi-view contrastive learning objective that promotes discriminability, semantic consistency, and structural coherence across intra- and inter-modal views. Extensive evaluations on two widely-used benchmark datasets, CMU-MOSI and CMU-MOSEI, demonstrate that SSU consistently achieves state-of-the-art performance while significantly reducing computational overhead compared to prior methods. Comprehensive qualitative analyses further validate SSU’s interpretability and its ability to capture nuanced emotional patterns through semantically-grounded interactions.
Jiangfeng Sun 0003, Sihao He 0001, Zhonghong Ou, Meina Song
AAAI4
2026 LDCS-Net: Local Dual-Context Attention Network for 3D Semantic Segmentation
abstract
High-precision 3D point cloud semantic segmentation in complex environments is often compromised by blurry boundaries and lost fine-grained details, primarily due to restricted receptive fields and geometric information decay. We present the Local Dual-Context and Multi-Scale Semantic Segmentation Network (LDCS-Net), a novel framework that preserves structural integrity through attention-guided graph convolution and dual-context fusion. By explicitly decoupling geometric and semantic cues, LDCS-Net strengthens local feature encoding and integrates dynamic graph construction with global attention to model long-range dependencies. Extensive experiments on the S3DIS dataset demonstrate that our method improves boundary delineation and small object recognition, achieving 65.9% mIoU and 72.2% mAcc. This represents a 1.2% boost in mIoU over the current state-of-the-art and a 3.2% increase in mAcc over the previous best-performing baseline. The source code is available at https://github.com/666-cute-yang/LDCS-NET.
Ying Zhou 0017, Meina Song, Huilin Ai, Jialong Li 0006, Zi Yang Chen
ICMR3
2026 Exploring Inter-Domain Wasserstein Metric for Adaptive Object Detection
abstract
Cross-domain adaptation has achieved significant development in recent years. Nevertheless, the models' performance varies dramatically across different scenarios. How to quantitatively measure inter-domain discrepancies to guide model training remains a challenging problem. Existing methods mainly focus on the trivial pixel-wise differences between the cross domain images, while they ignore the holistic discrepancies of the domain-specific distributions in various scenarios. Thus their effectivenesses are greatly limited in realistic applications. In this paper, we propose a universal method for measuring inter domain discrepancies based on Wasserstein distance. It alleviates the impact of intra-domain discrepancies on measurements and enables precise and quantitative representation of inter-domain discrepancies. We further integrate this metric into the generative models, and propose an assistant domain to conduct domain knowledge transfer for cross-domain object detection task. Experiments on three benchmarks validate the effectiveness of the proposed inter-domain measurement metric. Specifically, we achieve 51.8% mAP on CityScapes, 45.3% mAP on Clipart and 58.9% mAP on Watercolor, which are 1.5%, 0.5% and 0.8% higher than the state-of-the-art schemes, respectively.
Yanlong Lin, Ziqian Zhu, Yitian Guo, Zhonghong Ou, Siyuan Yao, Meina Song
IEEE Trans. Multim.6
2025 TSVC: Tripartite Learning with Semantic Variation Consistency for Robust Image-Text Retrieval
abstract
Cross-modal retrieval maps data under different modalities via semantic relevance. Existing approaches implicitly assume that data pairs are well-aligned and ignore the widely existing annotation noise, i.e., noisy correspondence (NC). Consequently, it inevitably causes performance degradation. Despite attempts that employ the co-teaching paradigm with identical architectures to provide distinct data perspectives, the differences between these architectures primarily stem from random initialization. Thus, the model becomes increasingly homogeneous along with the training process. Consequently, the additional information brought by this paradigm is severely limited. In order to resolve this problem, we introduce Tripartite Learning with Semantic Variation Consistency (TSVC) for robust image-text retrieval. We design a tripartite cooperative learning mechanism comprising a Coordinator, a Master, and an Assistant model. The Coordinator distributes data, and the Assistant model supports the Master model's noisy label prediction with diverse data. Moreover, we introduce a soft label estimation method based on mutual information variation, which quantifies the noise in new samples and assigns corresponding soft labels. We also present a new loss function to enhance robustness and optimize training effectiveness. Extensive experiments on three widely used datasets demonstrate that, even at increasing noise ratios, TSVC exhibits significant advantages in retrieval accuracy and maintains stable training performance.
Shuai Lyu, Zijing Tian, Zhonghong Ou, Yifan Zhu 0001, Qiankun Ha, Haoran Luo 0001, Meina Song
AAAI8
2025 Complex Numerical Reasoning with Numerical Semantic Pre-training Framework
abstract
Multi-hop complex reasoning over incomplete knowledge graphs (KGs) has been extensively studied, but research on numerical knowledge graphs (NKGs) remains relatively limited.Recent approaches focus on separately encoding entities and numerical values, using neural networks to process query encodings for reasoning.However, in complex multi-hop reasoning tasks, numerical values are not merely symbols, and they carry specific semantics and logical relationships that must be accurately represented.In this work, we propose a Complex Numerical Reasoning with Numerical Semantic Pre-training Framework (CNR-NST).The CNR-NST framework can perform binary operations on numerical attributes in NKGs, enabling it to infer new numerical attributes from existing knowledge.Our approach effectively handles up to 102 types of complex numerical reasoning queries.On three public datasets, CNR-NST demonstrates SOTA performance in complex numerical queries, achieving an average improvement of over 40% compared to existing methods.Notably, this work expands the query types for complex multi-hop numerical reasoning and introduces a new evaluation metric for numerical answers, which has been validated through comprehensive experiments.
Haihong E, Yifan Zhu 0001, Meina Song, Haoran Luo 0001
EMNLP5
2025 INFER: A Neural-symbolic Model For Extrapolation Reasoning on Temporal Knowledge Graph
abstract
Temporal Knowledge Graph(TKG) serves as an efficacious way to store dynamic facts in real-world. Extrapolation reasoning on TKGs, which aims at predicting possible future events, has attracted consistent research interest. Recently, some rule-based methods have been proposed, which are considered more interpretable compared with embedding-based methods. Existing rule-based methods apply rules through path matching or subgraph extraction, which falls short in inference ability and suffers from missing facts in TKGs. Besides, during rule application period, these methods consider the standing of facts as a binary 0 or 1 problem and ignores the validity as well as frequency of historical facts under temporal settings. In this paper, by designing a novel paradigm for rule application, we propose INFER, a neural-symbolic model for TKG extrapolation. With the introduction of Temporal Validity Function, INFER firstly considers the frequency and validity of historical facts and extends the truth value of facts into continuous real number to better adapt for temporal settings. INFER builds Temporal Weight Matrices with a pre-trained static KG embedding model to enhance its inference ability. Moreover, to facilitates potential integration with existing embedding-based methods, INFER adopts a rule projection module which enables it apply rules through conducting matrices operation on GPU. This feature also improves the efficiency of rule application. Experimental results show that INFER achieves state-of-the-art performance on various TKG datasets and significantly outperforms existing rule-based models on our modified, more sparse TKG datasets, which demonstrates the superiority of our model in inference ability.
Ningyuan Li 0002, Haihong E, Tianyu Yao, Haoran Luo 0001, Meina Song, Yifan Zhu 0001
ICLR7
2025 KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search
abstract
Knowledge Base Question Answering (KBQA) aims to answer natural language questions with a large-scale structured knowledge base (KB). Despite advancements with large language models (LLMs), KBQA still faces challenges in weak KB awareness, imbalance between effectiveness and efficiency, and high reliance on annotated data. To address these challenges, we propose KBQA-o1, a novel agentic KBQA method with Monte Carlo Tree Search (MCTS). It introduces a ReAct-based agent process for stepwise logical form generation with KB environment exploration. Moreover, it employs MCTS, a heuristic search method driven by policy and reward models, to balance agentic exploration’s performance and search space. With heuristic exploration, KBQA-o1 generates high-quality annotations for further improvement by incremental fine-tuning. Experimental results show that KBQA-o1 outperforms previous low-resource KBQA methods with limited annotated data, boosting Llama-3.1-8B model’s GrailQA F1 performance to 78.5% compared to 48.5% of the previous sota method with GPT-3.5-turbo. Our code is publicly available.
Haoran Luo 0001, Haihong E, Yikai Guo, Qika Lin, Xiaobao Wu, Xinyu Mu, Meina Song, Yifan Zhu 0001, Anh Tuan Luu
ICML8
2025 Towards Recognizing Spatial-temporal Collaboration of EEG Phase Brain Networks for Emotion Understanding
abstract
Emotion recognition from EEG signals is crucial for understanding complex brain dynamics. Existing methods typically rely on static frequency bands and graph convolutional networks (GCNs) to model brain connectivity. However, EEG signals are inherently non-stationary and exhibit substantial individual variability, making static-band approaches inadequate for capturing their dynamic properties. Moreover, spatial-temporal dependencies in EEG often lead to feature degradation during node aggregation, ultimately limiting recognition performance. To address these challenges, we propose the Spatial-Temporal Electroencephalograph Collaboration framework (Stella). Our approach introduces an Adaptive Bands Selection module (ABS) that dynamically extracts low- and high-frequency components, generating dual-path features comprising phase brain networks for connectivity modeling and time-series representations for local dynamics. To further mitigate feature degradation, the Fourier Graph Operator (FGO) operates in the spectral domain, while the Spatial-Temporal Encoder (STE) enhances representation stability and density. Extensive experiments on benchmark EEG datasets demonstrate that Stella achieves state-of-the-art performance in emotion recognition, offering valuable insights for graph-based modeling of non-stationary neural signals. The code is available at https://github.com/sun2017bupt/EEGBrainNetwork.
Jiangfeng Sun 0003, Kaiwen Xue 0001, Qika Lin, Yufei Qiao, Yifan Zhu 0001, Zhonghong Ou, Meina Song
IJCAI7
2025 DGFSD: Bridging the Gap between Dense and Sparse for Fully Sparse 3D Object Detection
abstract
Recently, LiDAR-based fully sparse 3D object detection has gained great attention, which utilizes point clouds to boost efficiency. Nevertheless, the relationship between well-studied dense representation and fully sparse representation is under-explored in existing studies, which focuses solely on building sparse representation by feature diffusion to solve the notorious center point missing problem. To this end, we propose a dense-guided fully sparse detection scheme, named DGFSD, to bridge the gap between dense and sparse features by dense-guided diffusion. Different from prior studies, we propose DgD (Dense-guided Diffusion) to overcome the center feature missing problem by dense knowledge transferring. Specifically, DgD transfers high-quality central point features from dense representations to endow sparse representations with dense knowledge. Moreover, we customize DFW (Dense Feature Weighting) to express uninformative representation and lift foreground representation. It makes high-quality dense feature contribute more to arcuate regression. To the best of our knowledge, we are the first to explore dense knowledge's impact on fully sparse framework. Extensive experiments conducted on nuScenes and Argoverse2 benchmark demonstrate the effectiveness of the proposed method. Specifically, DGFSD achieves 71.6% NDS and 67.3% mAP on the nuScenes test benchmark. On Argoverse2, DGFSD achieves 40.6% mAP, outperforming previous best hybrid and fully sparse methods. The code is available at https://github.com/Raiden-cn/DGFSD.
Zhonghong Ou, Kaiwen Xue 0001, Jiangfeng Sun 0003, Yifan Zhu 0001, Siyuan Yao, Yiran Shen 0007, Meina Song
ACM Multimedia8
2025 HyperGraphRAG: Retrieval-Augmented Generation via Hypergraph-Structured Knowledge Representation
abstract
Standard Retrieval-Augmented Generation (RAG) relies on chunk-based retrieval, whereas GraphRAG advances this approach by graph-based knowledge representation. However, existing graph-based RAG approaches are constrained by binary relations, as each edge in an ordinary graph connects only two entities, limiting their ability to represent the n-ary relations (n >= 2) in real-world knowledge. In this work, we propose HyperGraphRAG, the first hypergraph-based RAG method that represents n-ary relational facts via hyperedges. HyperGraphRAG consists of a comprehensive pipeline, including knowledge hypergraph construction, retrieval, and generation. Experiments across medicine, agriculture, computer science, and law demonstrate that HyperGraphRAG outperforms both standard RAG and previous graph-based RAG methods in answer accuracy, retrieval efficiency, and generation quality.
Haoran Luo 0001, Haihong E, Guanting Chen 0004, Yandan Zheng, Xiaobao Wu, Yikai Guo, Qika Lin, Yu Feng 0015, Zemin Kuang, Meina Song, Yifan Zhu 0001, Anh Tuan Luu
NeurIPS10
2025 PM-MOE: Mixture of Experts on Private Model Parameters for Personalized Federated Learning
abstract
Federated learning (FL) has gained widespread attention for its privacy-preserving and collaborative learning capabilities. Due to significant statistical heterogeneity, traditional FL struggles to generalize a shared model across diverse data domains. Personalized federated learning addresses this issue by dividing the model into a globally shared part and a locally private part, with the local model correcting representation biases introduced by the global model. Nevertheless, locally converged parameters more accurately capture domain-specific knowledge, and current methods overlook the potential benefits of these parameters. To address these limitations, we propose PM-MoE architecture. This architecture integrates a mixture of personalized modules and an energy-based personalized modules denoising, enabling each client to select beneficial personalized parameters from other clients. We applied the PM-MoE architecture to nine recent model-split-based personalized federated learning algorithms, achieving performance improvements with minimal additional training. Extensive experiments on six widely adopted datasets and two heterogeneity settings validate the effectiveness of our approach. The source code is available at https://github.com/dannis97500/PM-MOE.
Yu Feng 0015, Yifan Zhu 0001, Zongfu Han, Xie Yu, Kaiwen Xue 0001, Haoran Luo 0001, Mengyang Sun, Guangwei Zhang 0003, Meina Song
WWW10
2025 Multi-SEA: Multi-stage Semantic Enhancement and Aggregation for image-text retrieval
Zijing Tian, Zhonghong Ou, Yifan Zhu 0001, Shuai Lyu, Meina Song
Inf. Process. Manag.7
2025 An Improved Hybrid GC-LSTM Framework for Hourly Nowcasting of Ground-Level NO2 Concentrations Over Beijing-Tianjin- Hebei Region
abstract
Nitrogen dioxide (NO2) is a critical air pollutant with significant health and environmental implications, particularly in urban areas where high levels of emissions are prevalent. Accurate nowcasting of ground-level NO2 concentrations is essential for effective air quality management and timely public health interventions. Traditional methods often struggle with balancing the spatial accuracy of ensemble learning models and the temporal forecasting strengths of time-series models like long short-term memory (LSTM) networks. In this study, we propose an improved hybrid framework, GC-LSTM, to nowcast regional ground-level NO2 concentrations on an hourly scale based on satellite-derived NO2 vertical column densities (VCDs), meteorological data, and on-site observations. GC-LSTM integrates the spatial learning capabilities of grained cascade forest (gcForest) with the temporal prediction strengths of LSTM networks, leveraging the strengths of both spatial inference and time-series prediction. This study focuses on the Beijing-Tianjin–Hebei (BTH) region, one of China’s most polluted areas, as a case study. Our results indicate that the GC-LSTM framework performs a strong correlation between predicted and observed ground-level NO2 concentrations, with an$R^{2}$of 0.746 and a mean absolute percentage error (MAPE) of 18.4% at a 1-h prediction interval. Even as the prediction intervals extended to 2 and 3 h, the GC-LSTM consistently outperforms the gcForest model across all evaluated metrics, with$R^{2}$values higher by 0.097 and 0.117, and root mean square error (RMSE) values lower by 0.666 and$1.76~\mu \text {g/m}^{3}$than those nowcasted by using the standalone gcForest model, respectively, highlighting its robustness and adaptability. Furthermore, the capacity of the GC-LSTM framework for continual learning and adaptation ensures its effectiveness in dynamic environments, making it a valuable tool for real-time air quality forecasting and environmental management.
Zongfu Han, Meng Fan, Shipeng Song, Xiaoxia Liang, Meina Song, Guangyan He, Jinhua Tao, Liangfu Chen
IEEE Trans. Geosci. Remote. Sens.5
2025 Intradialytic Hypotension Frequency Prediction Using Generalizable Neighborhood Reasoning on Temporal Patient Knowledge Graph
abstract
Intradialytic hypotension (IDH) is a common complication among hemodialysis patients, adversely affecting quality of life and elevating mortality risk. IDH prediction enables physicians to take proactive measures, effectively reducing its occurrence. However, most prediction works rely on machine learning models, with a focus on real-time or session-level IDH. Hemodialysis patient data is multi-type and temporal, necessitating research on patient condition representation and temporal information utilization. Knowledge graphs (KGs) offer flexible data modeling and encompass rich structured information. This study represents patients using KGs and reason on graph structures to predict IDH. To study monthly IDH and utilize temporal information, a temporal patient KG is constructed. Patient KGs are first built at the monthly granularity based on data of 532 patients between January 2017 and August 2022. Six sequential monthly KGs are then combined into an observation window, resulting in a temporal KG dataset of 15,807 independent windows from 458 patients. The aim of this study is to utilize information from multiple months within a window to predict frequent IDH in the last month. However, the characteristics of IDH scenario and generalizability requirement pose challenges for the application of general KG reasoning models. Therefore, we adopt neighborhood-based KG reasoning and devise a visible feature guided patient-centric graph convolution to obtain patients' generalizable representations. Finally, patient representations in a window are fused using a sequential model, and processed by a prediction MLP to obtain the prediction results. Compared to 7 classic machine learning models, our model demonstrates superior performance in comprehensive metrics such as accuracy and F1 score.
Gengxian Zhou, Haihong E, Zemin Kuang, Tianyu Yao, Meina Song
IEEE J. Biomed. Health Informatics6
2024 CP-Prompt: Composition-Based Cross-modal Prompting for Domain-Incremental Continual Learning
Yu Feng 0015, Yifan Zhu 0001, Zongfu Han, Haoran Luo 0001, Guangwei Zhang 0003, Meina Song
ACM Multimedia7
2024 Text2NKG: Fine-Grained N-ary Relation Extraction for N-ary relational Knowledge Graph Construction
abstract
Beyond traditional binary relational facts, n-ary relational knowledge graphs (NKGs) are comprised of n-ary relational facts containing more than two entities, which are closer to real-world facts with broader applications. However, the construction of NKGs remains at a coarse-grained level, which is always in a single schema, ignoring the order and variable arity of entities. To address these restrictions, we propose Text2NKG, a novel fine-grained n-ary relation extraction framework for n-ary relational knowledge graph construction. We introduce a span-tuple classification approach with hetero-ordered merging and output merging to accomplish fine-grained n-ary relation extraction in different arity. Furthermore, Text2NKG supports four typical NKG schemas: hyper-relational schema, event-based schema, role-based schema, and hypergraph-based schema, with high flexibility and practicality. The experimental results demonstrate that Text2NKG achieves state-of-the-art performance in F1 scores on the fine-grained n-ary relation extraction benchmark. Our code and datasets are publicly available.
Haoran Luo 0001, Haihong E, Yuhao Yang 0006, Tianyu Yao, Yikai Guo, Zichen Tang, Wentai Zhang 0004, Shiyao Peng, Kaiyang Wan, Meina Song, Yifan Zhu 0001, Anh Tuan Luu
NeurIPS10
2024 GridMask: An Efficient Scheme for Real Time Curved Scene Text Detection
Zhonghong Ou, Siyuan Yao, Meina Song
PRCV (7)4
2024 SIIR: Symmetrical Information Interaction Modeling for News Recommendation
abstract
Accurate matching between user and candidate news plays a fundamental role in news recommendation. Most existing studies capture fine-grained user interests through effective user modeling. Nevertheless, user interest representations are often extracted from multiple history news items, while candidate news representations are learned from specific news items. The asymmetry of information density causes invalid matching of user interests and candidate news, which severely affects the click-through rate prediction for specific candidate news. To resolve the problems mentioned above, we propose a symmetrical information interaction modeling for news recommendation (SIIR) in this article. We first design a light interactive attention network for user (LIAU) modeling to extract user interests related to the candidate news and reduce interference of noise effectively. LIAU overcomes the shortcomings of complex structure and high training costs of conventional interaction-based models and makes full use of domain-specific interest tendencies of users. We then propose a novel heterogeneous graph neural network (HGNN) to enhance candidate news representation through the potential relations among news. HGNN builds a candidate news enhancement scheme without user interaction to further facilitate accurate matching with user interests, which mitigates the cold-start problem effectively. Experiments on two realistic news datasets, i.e., MIND and Adressa, demonstrate that SIIR outperforms the state-of-the-art (SOTA) single-model methods by a large margin.
Zhonghong Ou, Zongzhi Han, Peihang Liu, Shengyu Teng, Meina Song
IEEE Trans. Neural Networks Learn. Syst.5
2023 HAHE: Hierarchical Attention for Hyper-Relational Knowledge Graphs in Global and Local Level
abstract
Haoran Luo, Haihong E, Yuhao Yang, Yikai Guo, Mingzhi Sun, Tianyu Yao, Zichen Tang, Kaiyang Wan, Meina Song, Wei Lin. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Haoran Luo 0001, Haihong E, Yuhao Yang 0006, Yikai Guo, Mingzhi Sun, Tianyu Yao, Zichen Tang, Kaiyang Wan, Meina Song
ACL (1)9
2023 Hybrid Convolution Method for Graph Classification Using Hierarchical Topology Feature
Jiangfeng Sun 0003, Xinyue Lin, Fangyu Hao, Meina Song
ACML4
2023 LorenTzE: Temporal Knowledge Graph Embedding Based on Lorentz Transformation
Ningyuan Li 0002, Haihong E, Xueyuan Lin, Meina Song
ICANN (6)5
2023 FM-IGNN: Interaction Graph Neural Network with Fine-grained Matching for Session-based Recommendation
abstract
Session-based recommendation plays a critical role in a number of scenarios, e.g., e-commerce, which predicts user behavior based on sessions. The primary challenge is to match user interests and candidate items accurately. Existing studies mainly learn the representation of session items and candidate items independently, and aggregate session representations through an attention mechanism. Nevertheless, user interests are usually cross-domain, and the characteristics of the items are also multifactorial. Fixed session and candidate item representations lead to suboptimal matching. To some extent, it limits the representation capability of the model. In this paper, we present an Interaction Graph Neural Network with Fine-grained Matching, named FM-IGNN, for session-based recommendation. It incorporates interactive features into item and session modeling to match candidate items and user interests more accurately. We first propose an Interaction Graph Neural Network(IGNN) to learn candidate-aware session item representation and session-aware candidate item representation interactively. We then design a Fine-grained interest Matching (FM) framework to learn user interests related to candidate items, and calculate the matching scores of users and candidate items at multi-levels. Experiments on three benchmark datasets demonstrate that FM-IGNN outperforms the state-of-the-art(SOTA) schemes with a large margin. Specifically, on the Tmall dataset, it achieves relative improvement of up to 20.54%, 16.07%, 13.41%, and 15.52% on the Precision@10, MRR@10, Precision@20, and MRR@20, respectively.
Zongzhi Han, Zhonghong Ou, Yifan Zhu 0001, Meina Song
ICDM5
2023 Augmentation-free Dense Contrastive Distillation for Efficient Semantic Segmentation
Meina Song, Anbang Yao
NeurIPS4
2023 AD-RCNN: Adaptive Dynamic Neural Network for Small Object Detection
abstract
With the large-scale commercialization of 5G networks, Internet of Things (IoT) applications keep on emerging in recent years. Real-time environmental awareness is an essential part of various IoT applications, e.g., self-driving vehicles. Object detection plays a fundamental role in real-time environmental awareness, which is responsible for acquiring valuable object information from the environment automatically. Despite of the fast progress for object detection in general, small object detection still faces challenges. Because of the restricted scales, small objects are only capable of generating relatively week features after multiple convolutional layers, thus causing low detection accuracy. Existing schemes mostly focus on extracting rich multiscale features, e.g., generating high-resolution features through generative adversarial networks (GANs), or generating multiscale features through feature combination. Nevertheless, these schemes require complex network implementation, and usually suffer from high processing delay because of high-resolution images. To resolve the problems mentioned above, we propose an adaptive dynamic neural network (AD-RCNN) that consists of three fundamental improvements. We first propose a dynamic region proposal network to improve the quality of region proposals. We then introduce a visual attention scheme to generate features of regions. Finally, we put forward an adaptive dynamic training module to optimize final detection results. Experimental results demonstrate that AD-RCNN outperforms the state-of-the-art from the perspectives of mAP and frames per second (FPS). Specifically, at the resolution of 1024 of TT100K data set, AD-RCNN achieves 68.8% mAP, which outperforms the baseline Faster RCNN by 8.52%.
Zhonghong Ou, Zhaofengnian Wang, Fenrui Xiao, Baiqiao Xiong, Meina Song, Zheng Yan 0002, Pan Hui 0001
IEEE Internet Things J.6
2023 A knowledge distilled attention-based latent information extraction network for sequential user behavior
Ruo Huang, Shelby McIntyre, Meina Song, Haihong E, Zhonghong Ou
Multim. Tools Appl.3
2023 Free$\rm ^{3}$Net: Gliding Free, Orientation Free, and Anchor Free Network for Oriented Object Detection
abstract
Object detection for aerial images has achieved remarkable progress in recent years. Nevertheless, most exiting studies do not differentiate oriented object detection from horizontal detection. Certain schemes ignore the ambiguity of oriented object representation and leverage label assignment designed for horizontal object detection directly. Consequently, it leads to unstable training and causes performance degradation, because high-quality samples surrounding the oriented bounding boxes can not be leveraged effectively. To address this problem, we propose a gliding Free, orientation Free, and anchor Free Network (Free$\rm ^{3}$Net) with high-efficiency for oriented object detection. Specifically, we propose an unambiguous oriented object representation scheme, named FreeGliding, by gliding the projection points of samples on each edge of horizontal bounding boxes. It makes the detection largely free from representation ambiguity and multi-task dependency. To overcome the restrictions of label assignment, we put forward a novel Loss-aware Outer Sample Selection (LOSS) scheme, which takes into consideration spatial information and localization capability to retain high-quality samples surrounding the objects. Moreover, we introduce an Oriented Feature Fusion (OFF) scheme to tackle feature alignment by adjusting the receptive field and fusing oriented features dynamically. Experimental results on two large-scale remote sensing datasets HRSC2016 and DOTA demonstrate that Free$\rm ^{3}$Net outperforms the state-of-the-art schemes with a large margin. We hope our work can inspire rethinking the design of anchor-free detectors, and serve as a strong baseline for oriented object detection.
Zhonghong Ou, Zhongjie Chen, Shengyi Shen, Lina Fan, Siyuan Yao, Meina Song, Pan Hui 0001
IEEE Trans. Multim.6
2022 TCVM: Temporal Contrasting Video Montage Framework for Self-supervised Video Representation Learning
Fengrui Tian, Xie Yu, Shaoyi Du, Meina Song
ACCV (2)5
2022 Episodic Projection Network for Out-of-Distribution Detection in Few-shot Learning
abstract
The increasing demands of safety-critical computer vision applications have attracted extensive research on Out-of-Distribution (OOD) detection in recent years. Nevertheless, a large proportion of real-world tasks are in low-data regime, and the gap between meta-learning paradigm and OOD detection mechanism causes low performance in few-shot settings. In order to bridge the gap, we first propose an simple yet effective Episodic Projection Scheme (EPS). EPS is designed to project feature vectors to task-specific feature space for OOD detection, without sacrificing generalization of few-shot models. We then construct a multi-modal representation space for few-shot OOD detection by employing representations of the labels and their synonyms. At last, we put forward a few-shot OOD detection framework named Episodic Projection Network (EPN), which can integrate many kinds of perturbation based OOD algorithms with ease. To verify effectiveness of the proposed scheme, we implement several OOD algorithms into EPN and conduct experiments on two few-shot classification datasets, i.e., Omniglot and mini-ImageNet. Experimental results demonstrate that accuracy has been increased by 5% by integrating the OOD algorithms into the EPN framework.
Zhonghong Ou, Xie Yu, Shigeng Wang, Xiaoyang Kang 0002, Meina Song
ICPR8
2021 RTFE: A Recursive Temporal Fact Embedding Framework for Temporal Knowledge Graph Completion
abstract
Youri Xu, Haihong E, Meina Song, Wenyu Song, Xiaodong Lv, Wang Haotian, Yang Jinrui. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Youri Xu, Haihong E, Meina Song, Wenyu Song, Haotian Wang 0004, Jinrui Yang
NAACL-HLT3
2021 Redundancy Removing Aggregation Network With Distance Calibration for Video Face Recognition
abstract
Attention-based techniques have been successfully used for rating image quality, and have been widely employed for set-based face recognition. Nevertheless, for video face recognition, where the base convolutional neural network (CNN) trained on large-scale data already provides discriminative features, fusing features with only predicted quality scores to generate representation are likely to cause duplicate sample dominant problem, and degrade performance correspondingly. To resolve the problem mentioned above, we propose a redundancy removing aggregation network (RRAN) for video face recognition. Compared with other quality-aware aggregation schemes, RRAN can take advantage of similarity information to tackle the noise introduced by redundant video frames. By leveraging metric learning, RRAN introduces a distance calibration scheme to align distance distributions of negative pairs of different video representations, which improves the accuracy under a uniform threshold. A series of experiments is conducted on multiple realistic data sets to evaluate the performance of RRAN, including YouTube Faces, IJB-A, and IJB-C. In comprehensive experiments, we demonstrate that our method can diminish the overall influence of poor quality components with large proportion in the video and further improve the overall recognition performance with individual difference. Specifically, RRAN achieves a 96.84% accuracy on YouTube Face, outperforming all existing aggregation schemes.
Zhonghong Ou, Meina Song, Zheng Yan 0002, Pan Hui 0001
IEEE Internet Things J.3
2020 KGAnet: a knowledge graph attention network for enhancing natural language inference
abstract
Abstract Natural language inference (NLI) is the basic task of many applications such as question answering and paraphrase recognition. Existing methods have solved the key issue of how the NLI model can benefit from external knowledge. Inspired by this, we attempt to further explore the following two problems: (1) how to make better use of external knowledge when the total amount of such knowledge is constant and (2) how to bring external knowledge to the NLI model more conveniently in the application scenario. In this paper, we propose a novel joint training framework that consists of a modified graph attention network, called the knowledge graph attention network, and an NLI model. We demonstrate that the proposed method outperforms the existing method which introduces external knowledge, and we improve the performance of multiple NLI models without additional external knowledge.
Meina Song, Haihong E
Neural Comput. Appl.1
2019 A Novel Bi-directional Interrelated Model for Joint Intent Detection and Slot Filling
abstract
A spoken language understanding (SLU) system includes two main tasks, slot filling (SF) and intent detection (ID).The joint model for the two tasks is becoming a tendency in SLU.But the bi-directional interrelated connections between the intent and slots are not established in the existing joint models.In this paper, we propose a novel bi-directional interrelated model for joint intent detection and slot filling.We introduce an SF-ID network to establish direct connections for the two tasks to help them promote each other mutually.Besides, we design an entirely new iteration mechanism inside the SF-ID network to enhance the bi-directional interrelated connections.The experimental results show that the relative improvement in the sentence-level semantic frame accuracy of our model is 3.79% and 5.42% on ATIS and Snips datasets, respectively, compared to the state-of-the-art model.
Haihong E, Peiqing Niu, Zhongfu Chen, Meina Song
ACL (1)4
2019 FPSeq: Simplifying and Accelerating Task-Oriented Dialogue Systems via Fully Parallel Sequence-to-Sequence Framework
abstract
A mainstream task-oriented dialogue system is stuck in the independence of modules since it follows pipeline design, and it also suffers from sequence dependence and time dependence of recurrent neural network (RNN). Thus, these systems are complicated and usually trained slowly. In this paper, we propose FPSeq, a novel, fully parallel framework to simplify and accelerate task-oriented dialogue systems. Specifically, FPSeq turns pipeline design into a single sequence-to-sequence (seq2seq) model to achieve integration of each module for end-to-end training. In addition, multi-layer convolutional neural networks (CNNs) and attention mechanisms are applied in seq2seq learning to achieve parallel computations for speed improvement. Compared to the existing best model, the training speed of FPSeq is 3-10 times faster while only one-third of the number of parameters. Experimental results on CamRest676 and KVRET datasets indicate that FPSeq achieves the state-of-the-art performance in both task completion and quality of language generation.
Meina Song, Zhongfu Chen, Peiqing Niu, Haihong E
ICTAI1
2018 Is cloud storage ready? Performance comparison of representative IP-based storage systems
Zhonghong Ou, Meina Song, Zhen-Huan Hwang, Antti Ylä-Jääski, Ren Wang 0001, Yong Cui 0001, Pan Hui 0001
J. Syst. Softw.2
2017 Research and Implementation of Question Classification Model in Q&A System
Haihong E, Yingxi Hu, Meina Song, Zhonghong Ou
ICA3PP3
2017 A CNN-Based Supermarket Auto-Counting System
Zhonghong Ou, Changwei Lin, Meina Song, Haihong E
ICA3PP3
2017 Context-aware probabilistic matrix factorization modeling for point-of-interest recommendation
Xingyi Ren, Meina Song, Haihong E, Junde Song
Neurocomputing2
2017 Statistics-based CRM approach via time series segmenting RFM on large scale data
Meina Song, Xuejun Zhao, Haihong E, Zhonghong Ou
Knowl. Based Syst.1
2017 A three-stage global optimization method for server selection in content delivery networks
Junde Song, Meina Song
Soft Comput.3
2017 Exploring Vision-Based Techniques for Outdoor Positioning Systems: A Feasibility Study
abstract
Recent advances from wearables have significantly changed the way how humans communicate with the surrounding environment. To some extent, they have extended and augmented the capability of humans. For example, with a Google Glass, people can take pictures simply by winking eyes twice, which releases human hands from the cumbersome image-taking process. Thus, it enables new application scenarios that were not possible before. In this paper, we investigate utilizing vision-based techniques to provide a wearable positioning system. Specifically, we propose a Human-centric Positioning System (HoPS) that utilizes traffic signposts together with context information for real-time positioning. Towards that direction, we make three primary contributions: (1) we make several important observations that guide our design of HoPS system; for example, we find out that approximately 40 percent of traffic signposts monopolize a cell tower, and there are at most six signposts within the coverage of a single cell tower; (2) we investigate the impact factors of object detection success rate, and find its correlation with image quality, and resolution; and (3) we design and implement HoPS and an advanced version of HoPS based on additional context information from Wi-Fi network, which we name HoPS-WiFi. Experimental results demonstrate the effectiveness of HoPS, especially HoPS-WiFi, which can estimate the relevant location correctly within 1.3 seconds.
Meina Song, Zhonghong Ou, Eduardo Castellanos, Tuomas Ylipiha, Teemu Kämäräinen, Matti Siekkinen, Antti Ylä-Jääski, Pan Hui 0001
IEEE Trans. Mob. Comput.1
2016 TGTM: Temporal-Geographical Topic Model for Point-of-Interest Recommendation
Cong Zheng, Haihong E, Meina Song, Junde Song
DASFAA (1)3
2016 CMPTF: Contextual Modeling Probabilistic Tensor Factorization for recommender systems
Cong Zheng, Haihong E, Meina Song, Junde Song
Neurocomputing3
2015 User Familiar Degree Aware Recommender System
abstract
In a recommender system, items can be rated across multiple fields by users with varying degrees of familiarity. Hence, the ratings in a recommender system should have different recommended weights. Ratings in fields where in the user has high or low familiarity should be given high or low recommended weights, respectively. However, current recommendation algorithms ignore this problem and use the ratings indiscriminately, thus affecting the accuracy of the recommendation system. In this paper, we provide a focused study of user-familiarity degree-aware recommendation and develop a user-familiarity degree-aware latent factor model for recommendations that considers both user familiarity and item features reflected by the tagging information. We also design a user-familiarity degree-aware probability matrix factorization model, which computes the degree of familiarity of a user with the items he/she has rated. By using the user-familiarity degree, different recommended weights are given to every rating to obtain precise recommendations. The experiment results on real-world datasets show that our algorithm significantly outperforms state-of-the-art latent factor models and effectively improves the accuracy of the recommendation results.
Yusheng Li 0005, Haihong E, Meina Song, Junde Song
ICWS3
2015 Cost-efficient coordinated scheduling for leasing cloud resources on hybrid workloads
Sen Su, Xiang Cheng 0003, Meina Song, Liyu Ma, Jie Wang 0002
Parallel Comput.4
2014 Hierarchical prediction based task scheduling in hybrid data center
abstract
Cloud computing can help data center consolidate batch and gratis tasks with over-provisioned production applications, and fulfill their diverse resource demands and performance objectives with high scalability and flexibility. One challenge in this hybrid data center is that the dramatic fluctuation of batch and gratis workload may impact performance of production applications, cause task failure, decrease efficiency, and waste computing resources. One way to tackle the challenge is to reduce resource allocation to prevent host overload by delay scheduling tasks if resources are predicted in short. In this paper, we propose hierarchical prediction method for hybrid workload. We use last-state based ARMA model to predict stationary process of production workload, and use feedback based online AR model to predict the vibrated workload of batch and gratis tasks. Evaluation shows that the hierarchical prediction based task scheduling can reduce host overload by more than 85 percent, reduce tasks evicted and killed by more than 60 percent, and reduce 40 percent of average task scheduling delay.
Haiou Jiang, Haihong E, Meina Song
ICPADS3
2014 Incorporating appraisal expression patterns into topic modeling for aspect and sentiment word identification
Kwei-Jay Lin, Meina Song
Knowl. Based Syst.5
2011 HEaRS: A Hierarchical Energy-Aware Resource Scheduler for Virtualized Data Centers
abstract
With the increasing popularity of Internet-based cloud services, energy efficiency in large-scale Internet data centers has become important not only to curtail energy costs and alleviate environmental concern, but also because such systems can quickly reach the limits of power available to them. This paper investigates to what extent and how energy usage improvements through consolidation can benefit from taking into account the environmental influences and effects seen in data center systems. Toward that end, we present experimental results obtained in a fully instrumented, small scale data center and then use these results to propose a hierarchical energy-aware resource scheduler (HEaRS) for cluster workload placement and server provisioning, also considers the physical environment in which data center systems operate. Specifically, at the rack level, HEaRS tries to maintain a 'thermal balance' across the rack to avoid hot spots and reduce cooling costs. At the chassis level, HEaRS utilizes the proportional plus integral controller to achieve a balance in the levels of usage of electrical current between the two power domains in the chassis, which helps the chassis reach its most energy efficient state. Finally, at server level, HEaRS can employ known methods like dynamic voltage and frequency scaling or core idling to reduce power consumption. This results in a hierarchical set of controllers that jointly, implement holistic solutions to energy-aware resource scheduling for an entire rack, and this hierarchical solution can then be further extended to entire data centers. Our initial experiment result show opportunities for gains, with up to 16% in energy usage compared to methods that are not aware of the physical environment and up to 15% improvements in application performance.
Meina Song, Junde Song, Ada Gavrilovska, Karsten Schwan
CLUSTER2
2010 The Research of Service Network Based on Complex Network
abstract
At present the service science and engineering research are mainly focus on service discovery, service composition, service reputation and other key technical of the service computing. However the basic theory of service science, in particular, the basic principle of service and the evolution mechanism of service network only have some preliminary of research and exploration. The atomic services as the network nodes and the relationships of services combination as the network edges constitute the service network based services provision environment. This paper researched the basic characteristics of services and service networks, and proposed a new research method to explore the service network's "small world", "scale-free" characteristics and service network topology, based on the theory of complex network and existing networked software research works. It presents a new research perspective and methods for the service science and engineering research.
Haihong E, Meina Song, Junde Song, Zhijun Ren
ICSS2
2010 QoS Oriented Cross-Layer Design for Supporting Multimedia Services in Cooperative Networks
abstract
Providing Quality of Service (QoS) for multimedia communication is a challenging problem in wireless multi-hop networks using cooperative diversity. In this paper, we present a fully distributed cross-layer design algorithm incorporating delay constraint in order to support real-time multimedia services. We utilize network utility maximization framework in the cross-layer design algorithm modeling. The proposed algorithm joint solve congestion-control problem and routing problem using convex optimization method. Moreover, we design an End-to-end Delay Framework for delay constraint modeling when analyzing the network utility maximization problem such that QoS guarantees for multimedia communication is enhanced. Convergence Analysis indicates that this algorithm can converge, which is the expected outcome for providing QoS in supporting multimedia services.
Meina Song
ICSS2
2010 Home Appliance Mashup System Based on Web Service
abstract
With the development of Network Appliance, more and more mechanisms about network control have been introduced to improve the performance and the user experience. Since the cost of web-based interfaces is considerably low, Web can be used to provide the infrastructure for the design of simple and user-friendly interfaces for household appliances. In this paper, we propose a new system which can share the home appliance abilities to Internet as Web Service. By our system, a lot of Web Based Network Appliance related applications can be developed easily and give end users more friendly, flexible control interfaces. Furthermore, functions of different appliances can be integrated to provide more complex service to user, which is named Home Appliance Mashup in this paper.
Meina Song
ICSS2
2010 A Mobile P2P Community System
abstract
The distributed architecture of vitual community based on P2P has proved to resolves some problems from central architecture such as single point failure. In this paper, drawing on the idea of SOA, we propose an optimized architecture of distributed mobile P2P community system, in which the virtual communities are divided into separately service components which can be deployed in the overlay networks as resource instead of a central server. And the introduction of the physical network aware technique of P4P can help the client find the most suitable service replica. We also propose a replica consistent control mechanism to facilitate the practice of the system in this paper. At last we develope a prototype system of mobile community.
Meina Song
ICSS2
2010 Generic service composition platform for pervasive E-Commerce
abstract
Abstract In this paper, a generic E‐Commerce process model with full E‐Commerce service coverage is propose. Based on this model, a generic service composition platform is proposed as an individual middleware layer to support the pervasive E‐Commerce. The platform consists of an orchestration component and several atomic service components (ASCs). It is in charge of dynamic and automatic general service (GS) construction taking the user preferences and demanded quality of service (QoS) into account. The designed platform is flexible and extensible thanks to the service oriented architecture (SOA) (service oriented architecture)‐based atomic reusable and sharable service components (SCs). And it employs efficient peer‐to‐peer (P2P) interaction between two components. In our method, global QoS is guaranteed by the runtime QoS monitoring and dynamic adjusted two‐dimensional orchestration plan. Copyright © 2009 John Wiley & Sons, Ltd.
Jing Chi, Chunyang Yin, Meina Song, Junde Song, Xiaosu Zhan
Wirel. Commun. Mob. Comput.3
2008 Extended WDB Algorithm for QoS Enhancement in IEEE 802.11e WLAN
abstract
The paper proposes an extended waiting-time dependent backoff (EWDB) algorithm for IEEE 802.11e WLAN as an extension of our previous WDB method. Concerning the QoS limitations of EDCA, EWDB controls backoff window (BW) according to waiting-time (WT), and staggers BW ranges among prioritized services. The algorithm is based on a flexible mathematic model consisting of deterministic function and random function. Backoff functions are respectively designed for voice, video and data. WT-dependent sub-priority is introduced for voice and video. An overlap-restricted backoff function is designed for data traffic. Simulation results show the superiority of EWDB to EDCA in terms of jitter and delay of voice/video, throughput of data, as well as overall collision rate under different network load.
Jing Chi, Meina Song, Junde Song
VTC Fall2