Yanni Ma

dblp:265/2454 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Edge-Centric Relational Reasoning for 3D Scene Graph Prediction
abstract
3D scene graph prediction aims to abstract complex 3D environments into structured graphs consisting of objects and their pairwise relationships. Existing approaches typically adopt object-centric graph neural networks, where relation edge features are iteratively updated by aggregating messages from connected object nodes. However, this design inherently restricts relation representations to pairwise object context, making it difficult to capture high-order relational dependencies that are essential for accurate relation prediction. To address this limitation, we propose a Link-guided Edge-centric relational reasoning framework with Object-aware fusion, namely LEO, which enables progressive reasoning from relation-level context to object-level understanding. Specifically, LEO first predicts potential links between object pairs to suppress irrelevant edges, and then transforms the original scene graph into a line graph where each relation is treated as a node. A line graph neural network is applied to perform edge-centric relational reasoning to capture inter-relation context. The enriched relation features are subsequently integrated into the original object-centric graph to enhance object-level reasoning and improve relation prediction. Our framework is model-agnostic and can be integrated with any existing object-centric method. Experiments on the 3DSSG dataset with two competitive baselines show consistent improvements, highlighting the effectiveness of our edge-to-object reasoning paradigm.
Yanni Ma, Hao Liu 0061, Yulan Guo, Theo Gevers, Martin R. Oswald
AAAI1
2026 WeJoy Mat: An Embodied Interactive Mat System Mediating Parent-Child Play for Families with Autistic Children
abstract
While interactive technologies show promise for autistic children, current designs often overlook the heterogeneity of this population, offering standardized experiences that fail to account for individual profiles. We present WeJoy Mat, a multi-sensory embodied system designed to mediate parent-child co-play through a mirrored game paradigm. We conducted a mixed-methods study with 11 dyads, triangulating gameplay logs, behavioral rating, and qualitative interviews. Our analysis revealed that interaction outcomes were not driven by specific sensory modalities alone, but were moderated by the fit between these modes and children’s individual cognitive-behavioral phenotypes. Consequently, the system functioned as an active "third actor" within the interaction triad, where parents provided scaffolding while the mat redistributed agency and facilitated parallel synchrony. Our work advocates for a shift from rigid presets to flexible, parent-controlled tools that effectively move design away from correction to truly neuro-affirming environments that support the unique needs of each child.
Jiayu Jiang, Shaolong Chai, Xinghao Jiang, Yanni Ma, Jingwei He, Fangtian Ying
IDC7
2025 Semi-Structured Interview System Based on Fine-Tuned Large Language Model and Reinforcement Learning from Human Feedback
abstract
Semi-structured interview is an important method in human-computer interaction research. However, traditional methods often rely on the experience and skills of the interviewer, limiting the quality and flexibility of interview outline. We propose a semi-structured interview system based on fine-tuned large language model (LLM) and reinforcement learning from human feedback (RLHF). Our system uses knowledge in the domain of human-computer interaction to fine-tune LLM and adopts the RLHF method based on multi-task learning and entropy regularization to dynamically adjust interview strategies. The system also collects multimodal data such as speech, video, and emotion to provide comprehensive support for subsequent analysis. Simulation experiment and pilot experiment show that compared to traditional human interview and other baseline methods, our system performs well in terms of interview relevance, personalization, engagement, and efficiency. Qualitative analysis further reveals the system's advantages in terms of conversational fluency, personalization, and efficiency. This innovative interview method is expected to play an important role in future user research, providing researchers with a more efficient and comprehensive data collection tool.
Yanni Ma, Shaolong Chai
CSCWD1
2025 A Bio-Inspired Design Method Based on LLM: A Case Study of Flying Car Concept Design
abstract
Bio-inspired design (BID), as an important method for product innovation, still has many challenges in knowledge retrieval efficiency, feature mapping and solution generation. We proposed a dialogical BID method based on large language model (LLM), aiming to enhance the efficiency and innovation of the design process through human-computer collaboration. Using the conceptual design of a flying car as a case study, we accomplished an innovative transformation from the biological features of a humpback whale to a product form through multiple rounds of dialog with LLM. Experimental evaluation showed that the method was outstanding in terms of novelty and fashion and symbolization. The study confirms the effectiveness of LLM in BID and provides new methodological ideas for product innovation.
Yanni Ma, Shaolong Chai
CSCWD2
2025 Understanding and Supporting Multimodal AI Chat Interactions of DHH College Students: an Empirical Study
Nan Zhuang, Yanni Ma, Shaolong Chai, Shitong Weng, Mengru Xue, Yuxi Mao
ICMI2
2024 Heterogeneous Graph Learning for Scene Graph Prediction in 3D Point Clouds
Yanni Ma, Hao Liu 0061, Yun Pei 0001, Yulan Guo
ECCV (26)1
2024 Spatial-Temporal Hazard Prediction of Rainfall-Induced Landslides Using Multi-Modal Earth Observation Data
abstract
Rainfall is the primary landslide triggering factor in China, and the spatial-temporal hazard prediction of rainfall-induced landslides is of great practical significance. Currently, most countries and regions establish landslide hazard prediction systems based on rainfall data only, resulting in low spatial precision of hazard prediction results and a high false alarm rate. This paper proposes a hazard prediction model that considers landslide triggering factors, landslide predisposing environment, and the spatial regularity of historical landslides based on multi-modal earth observation data. The proposed model has significantly improved the spatial-temporal hazard prediction performance of rainfall-induced natural terrain landslides in Hong Kong.
Yangyang Chen 0004, Junchuan Yu, Dongping Ming, Yanni Ma, Yuanbiao Dong, Rongyuan Liu, Daqing Ge
IGARSS5
2024 Quantification of Potential Ice Road Evolution in the Pan-Arctic and its Impacts Through Remote Sensing Observations
abstract
Ice roads serve as vital land transportation during the Arctic winter season. In the context of polar increased warming, there are great uncertainties for human activities on ice roads. In this paper, we integrate remote sensing techniques to quantify the potential ice road evolution in the Pan-Arctic and its impact on land accessibility from 1979 to 2017. We show that the potential ice roads have significantly decreased, with the fastest decrease in March to 2.34×104km2yr−1. Furthermore, the contribution of potential ice roads to port accessibility is most severely reduced in the Canadian Arctic, reaching 0.93 h yr−1. The results demonstrate that warmer winters are imposing severe stress on Arctic land access.
Yuanbiao Dong, Pengfeng Xiao, Daqing Ge, Junchuan Yu, Yangyang Chen 0004, Yanni Ma, Rongyuan Liu
IGARSS8
2024 The Extraction of Deformation Zone in Insar Based on the Lightweight Designed Model Bisnet Network
abstract
In the identification of geological hazards in a wide area, the rapid extraction of deformation areas in the InSAR phase becomes an important part of whether the hazards can be quickly and accurately identified. In practical applications, there is a significant difference in the size of deformation regions over a wide area. Therefore, this article uses the Bilateral Segmentation Network (Bisenet), which calculates in parallel through two branches. We first utilize a small step spatial path to preserve spatial information and generate high-resolution features. Meanwhile, a context path with a fast downsampling strategy is employed to obtain sufficient receptive fields. On top of these two paths, we utilize a new feature fusion module to effectively combine features. This greatly improves the extraction speed while preserving multilayer feature fusion and preserving information extraction results at different scales. A suitable balance has been achieved between speed and segmentation performance, which well meets the current work of identifying geological hazards in wide areas.
Yanni Ma, Yangyang Chen 0004, Junchuan Yu, Yuanbiao Dong
IGARSS1
2024 Comparison of Pixel-Level and Feature-Level Image Fusion Networks for Slow-Moving Landslide Detection
abstract
Slow-moving landslide detection is of vital importance in preventing and mitigating geohazards. Extracting abstract features from remote sensing images is crucial for achieving high-precision detection of slow-moving landslides. This study utilizes both activity features and terrain structure features for geohazard detection. We propose a pixel-level and a feature-level image fusion network, and investigate the multi-level fusion cooperative mechanism. We evaluate the performance of the two-level fusion and single-modal data base on the test data. The experimental results demonstrate that fusion of the activity characteristics and topographic characteristics can enhance the accuracy of identifying slow-moving landslides. Feature-level fusion outperforms pixel-level fusion for slow-moving landslides identification.
Yanni Ma, Yangyang Chen 0004, Yuanbiao Dong, Junchuan Yu, Daqing Ge
IGARSS2
2024 Landslidenet: Adaptive Vision Foundation Model for Landslide Detection
abstract
Recent advancements in Vison Foundation Models (VFMs) like the Segment Anything Model (SAM) have exhibited remarkable progress in natural image segmentation. However, its performance on remote sensing images is limited, especially in some application scenarios that require strong expert knowledge involvement, such as landslide detection. In this study, we proposed an effective segmentation model, namely LandslideNet, which is realized by embedding a tuning layer in a pre-trained encoder and adapting the SAM to the landslide detection scene for the first time. The proposed method is compared with traditional convolutional neural networks (CNN) on two well-known landslide datasets. The results indicate that the proposed model with fewer training parameters has better performance in detecting small-scale targets and delineating landslide boundaries, with an improvement of 6-7 percentage points in accuracy (F1 and mIoU) compared to mainstream CNN-based methods.
Junchuan Yu, Yichuan Li 0006, Yangyang Chen 0004, Changhong Hou, Daqing Ge, Yanni Ma
IGARSS6
2024 Hidden Scars: Anti-bullying Serious Game Design for Rural Children
Shaolong Chai, Yanni Ma, Fengyan Hu
ICEC4
2023 AnchorPoint: Query Design for Transformer-Based 3D Object Detection and Tracking
abstract
With the success of Transformers in natural language processing, object detection with Transformers (DETR) has attracted widespread attentions. In previous Transformer-based 2D detectors, the object queries are a set of learning embeddings. However, it is very hard to apply these detectors to the 3D domain due to the lack of explicit physical meanings and position priors of learned object queries. In this paper, we introduce the concept of anchors and propose a novel query design based on anchor points. In our query design, we use the foreground points as the anchor points and encode these anchor points as the object queries. Consequently, each object query has an explicit physical meaning and only focus on its nearby object. Additionally, we also propose an instance-aware sampling strategy to select a small set of representation foreground points from the scene point cloud. Extensive experiments on several large-scale 3D object detection datasets demonstrate that the proposed AnchorPoint detector achieves promising accuracy and efficiency. In particularly, AnchorPoint achieves an average precision (AP) of 83.21 at 61 frame-per-second (FPS) on the moderate level of the KITTI-DET Car subset. Moreover, we model each object as its corresponding anchor point, and extend the AnchorPoint model to 3D multi-object tracking by adding an extra tracking head. We show that our method achieves comparable performance to existing state-of-the-art methods on the KITTI-MOT dataset.
Hao Liu 0061, Yanni Ma, Hanyun Wang, Yulan Guo
IEEE Trans. Intell. Transp. Syst.2
2023 CenterTube: Tracking Multiple 3D Objects With 4D Tubelets in Dynamic Point Clouds
abstract
3D Multi-Object Tracking (MOT) in dynamic point cloud sequences is a fundamental research problem for several downstream tasks such as motion planning and action recognition. Existing methods usually rely on the traditional tracking-by-detection (TBD) paradigm, which performs the tracking based on the results achieved by dedicated detectors. However, this two-stage framework usually cannot sufficiently exploit spatial-temporal information and end-to-end optimization, leading to sub-optimal tracking performance, especially when the object is partially or completely occluded. In this paper, we propose a joint detection and tracking framework namedCenterTubefor dynamic point cloud sequences. The key to our approach is to formulate the problem of multiple object trajectory predictions as 4D tubelet detections. In particular, the proposed CenterTube is composed of three head branches, including a center branch, a regression branch, and a movement branch for the estimation of object center, object size, instance movement, and frame interval, respectively. Additionally, a Tube BEV-IoU (TB-IoU) is also presented to link the generated clip-level tubelets and form the final tracks. Extensive experiments conducted on the KITTI-MOT and nuScenes datasets demonstrate that our model achieves competitive performances even if no ready-made detection results is adopted.
Hao Liu 0061, Yanni Ma, Qingyong Hu, Yulan Guo
IEEE Trans. Multim.2
2021 High-performance pipeline architecture for packet classification accelerator in DPU
abstract
Packet classification is a fundamental problem in the network. With the rapid growth of network bandwidth, wire-speed packet classification has become a key challenge for next-generation network processors. In this paper, we propose a decision-tree-based, multi-pipeline architecture for packet classification accelerator in Data Processing Unit (DPU). Our solution is based on MBitTree, a memory-efficient decision tree algorithm for packet classification. First, we present a parallel architecture composed of multiple linear pipelines for efficiently mapping the decision tree built by MBitTree. Second, a special logic is designed to quickly traverse the decision tree, reducing the logic delay of the pipeline stage. Finally, several pipeline optimization techniques are proposed to improve the performance of the architecture. The implementation results show that our architecture can achieve more than 250 Gbps throughput for the 64-byte minimum Ethernet packets, and can store 100K rules in the on-chip memory of a single NetFPGA_SUME.
Gaofeng Lv, Yanni Ma, Guanjie Qiao
FPT3
2021 Semantic Context Encoding for Accurate 3D Point Cloud Segmentation
abstract
Semantic context plays a significant role in image segmentation. However, few prior works have explored semantic contexts for 3D point cloud segmentation. In this paper, we propose a simple yet effective Point Context Encoding (PointCE) module to capture semantic contexts of a point cloud and adaptively highlight intermediate feature maps. We also introduce a Semantic Context Encoding loss (SCE-loss) to supervise the network to learn rich semantic context features. To avoid hyperparameter tuning and achieve better convergence performance, we further propose a geometric mean loss to integrate both SCE-loss and segmentation loss. Our PointCE module is general and lightweight, and can be integrated into any point cloud segmentation architecture to improve its segmentation performance with only marginal extra overheads. Experimental results on the ScanNet, S3DIS and Semantic3D datasets show that consistent and significant improvement can be achieved for several different networks by integrating our PointCE module.
Hao Liu 0061, Yulan Guo, Yanni Ma, Yinjie Lei, GongJian Wen
IEEE Trans. Multim.3
2020 Global Context Reasoning for Semantic Segmentation of 3D Point Clouds
abstract
Global contextual dependency is important for semantic segmentation of 3D point clouds. However, most existing approaches stack feature extraction layers to enlarge the receptive field to aggregate more contextual information of points along the spatial dimension. In this paper, we propose a Point Global Context Reasoning (PointGCR) module to capture global contextual information along the channel dimension. In PointGCR, an undirected graph representation (namely, ChannelGraph) is used to learn channel independencies. Specifically, channel maps are first represented as graph nodes and the independencies between nodes are then represented as graph edges. PointGCR is a plug-andplay and end-to-end trainable module. It can easily be integrated into an existing segmentation network and achieves a significant performance improvement. We conduct extensive experiments to evaluate the proposed PointGCR module on both indoor and outdoor datasets. Experimental results show that our PointGCR module efficiently captures global contextual dependencies and significantly improve the segmentation performance of several existing networks.
Yanni Ma, Yulan Guo, Hao Liu 0061, Yinjie Lei, GongJian Wen
WACV1