VLDB 2026 Research / reviewers in the wild / expert
Yushan Pan
dblp:117/0714
· DBLP profile ↗
54ranked-venue papers
5as first author
51since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 1 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 12 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 9 since 2021Systems, architecture and hardware · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | QDRF: Quantized and distilled robust fusion for incomplete data in multimodal sentiment analysis
Jiachen Hou, Zheyan Cheng, Chen Xing, Yushi Li, Shaoyu Cai, Yushan Pan |
Inf. Sci. | 8 |
| 2026 | SEVADE: Self-Evolving Multi-Agent Analysis with Decoupled Evaluation for Hallucination-Resistant Sarcasm DetectionabstractSarcasm detection is a crucial yet challenging Natural Language Processing task. Existing Large Language Model methods are often limited by single-perspective analysis, static reasoning pathways, and a susceptibility to hallucination when processing complex ironic rhetoric, which impacts their accuracy and reliability. To address these challenges, we propose SEVADE, a novel Self-Evolving multi-agent Analysis framework with Decoupled Evaluation for hallucination-resistant sarcasm detection. The core of our framework is a Dynamic Agentive Reasoning Engine (DARE), which utilizes a team of specialized agents grounded in linguistic theory to perform a multifaceted deconstruction of the text and generate a structured reasoning chain. Subsequently, a separate lightweight rationale adjudicator (RA) performs the final classification based solely on this reasoning chain. This decoupled architecture is designed to mitigate the risk of hallucination by separating complex reasoning from the final judgment. Extensive experiments on four benchmark datasets demonstrate that our framework achieves state-of-the-art performance, with average improvements of 7.01% in Accuracy and 6.55% in Macro-F1 score. Mingxuan Hu, Yushan Pan, Zhijie Xu, Yangbin Chen |
AAAI | 5 |
| 2026 | Intrinsic vs. Extrinsic Programming Challenges in Educational Games: How they shape Children's Computational Thinking, Learning Drive, and Game EngagementabstractEducational programming games (EPGs) build computational thinking (CT), a vital 21st-century skill. A core design challenge is how pedagogical and gameplay challenges are integrated to balance educational objectives with player engagement. This study formalizes two contrasting challenge design patterns that reflect distinct integration strategies: extrinsic programming challenges (C1), a pedagogy-oriented design where programming is the core challenge enforced through external constraints; and intrinsic programming challenges (C2), a gameplay-oriented design where programming serves as a tool for overcoming gameplay challenges raised by in-game puzzles. To examine these challenge design patterns, we developed two isomorphic EPGs WannaBone1 (C1) and WannaBone2 (C2), each featuring 20 levels introducing sequences, loops, conditionals, and global variables. A controlled classroom study with 306 primary school students reveals that both designs improve CT, whereas C2 yields significantly higher intrinsic learning motivation and high-order immersion of flow. These findings indicate that a gameplay-oriented rather than pedagogy-oriented design perspective better unites education and entertainment, guiding future EPGs design. Baijun Chen, Yushan Pan, Shenze Shandel Huang |
CHI | 2 |
| 2026 | CircuitS2L: Circuit Dataset Augmentation via Generative Featuring and Supervised Labeling
Linyu Zhu, Tsun-Ming Tseng, Yushan Pan, Xinfei Guo |
ISCAS | 3 |
| 2026 | Hallux valgus diagnosis on X-rays with skeleton-guided diffusionabstractAbstract Medical image synthesis plays a crucial role in providing anatomically accurate images for diagnosis and treatment. Hallux valgus (HV), which affects approximately 19% of the global population, often requires frequent weight-bearing X-rays for assessment, which can be costly and logistically challenging in resource-limited healthcare settings due to the scarcity of horizontal foot imaging systems. Existing X-ray models often struggle to balance image fidelity, skeletal consistency, and physical constraints, particularly in diffusion-based methods that lack skeletal guidance. We propose the skeletal-constrained conditional diffusion model (SCCDM) along with two evaluation metrics, keypoint confidence-completeness (KCC) and keypoint structural consistency (KSC), to assess the anatomical plausibility of generated X-rays based on keypoint confidence and structural consistency. SCCDM incorporates multi-scale feature extraction and attention mechanisms, improving the structural similarity index by 5.72% (0.794) and the peak signal-to-noise ratio by 18.34% (21.40 dB). When combined with KCC and KSC, this method achieves average scores of 0.85 and 0.84, respectively, facilitating the clinical assessment of HV deformity and enabling more accurate surgical intervention. The code is available at https://github.com/midisec/SCCDM. Midi Wan, Yizhuo Liang 0002, Di Wu 0035, Yushan Pan, Guangzhen Zhu, Feng Qu, Hao Wang 0003 |
Comput. J. | 5 |
| 2026 | CR-GAC: Cross-modal Recombination via Graph-Attention Collaborative Optimization for multimodal sentiment analysis
Haoran Chen 0004, Zuhe Li, Yushan Pan, Hongwei Tao, Huaiguang Wu, Yunyang Wang, Chenguang Yang 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Algorithm-hardware co-design of binary neural network for efficient super resolution on FPGA
Yuanxin Su, Yushan Pan, Zhijie Xu, Xinfei Guo |
Integr. | 3 |
| 2026 | Bridging the Gap: Understanding Collaborative Practices of Underrepresented Student Developers in the GitHub CommunityabstractAbstract While studying student developers who engage with the GitHub community is not new, there is a notable lack of understanding regarding how underrepresented groups participate in collaborative efforts. Previous research has typically focused on the initial moment of joining summer programs rather than examining the entire contribution process over time. This approach limits insights into how motivations and challenges evolve during contributions and often centers on individual experiences rather than collaborative behaviors among multiple participants. Our 13-month online ethnographic study investigates the collaborative practices of student developers within an Apache open-source software (OSS) community on GitHub. We analyze their interactions using GitHub features and the impact of Apache guidelines. By exploring the achievements and challenges students face in coding activities, we provide actionable insights to bridge the gap between community expectations and student needs. Our findings highlight strategies such as targeted mentorship and improved communication to support sustained contributions, offering valuable guidance for fostering a more inclusive OSS ecosystem. Jingxiao Sun, Yushan Pan |
Interact. Comput. | 6 |
| 2026 | Innovative application of large language model in the prediction of potential depression in the elderly: based on CHARLS dataset
Minyan Su, Yushan Pan |
Multim. Syst. | 5 |
| 2026 | A text-guided cross-hierarchical fusion and multi-task learning framework for multimodal sentiment analysis
Yushan Pan, Zuhe Li, Di Wu 0035, Zhiyang Zhao, Yuanping Xu, Zhijie Xu |
Neural Networks | 2 |
| 2026 | AL-HCL: Active Learning and Hierarchical Contrastive Learning for Multimodal Sentiment Analysis With Fusion Guidance
Xiaojiang He, Yushan Pan, Zhijie Xu, Zuhe Li, Xinfei Guo, Chenguang Yang 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2026 | Clue and Context Fusion for Sarcasm Detection with Large Multimodal ModelsabstractDetecting sarcasm in social media is fundamentally different from general VLM benchmarks: it is a pragmatic contradiction problem in which the literal signal in one modality is intentionally misaligned with the intended meaning, while dominant pre-training (e.g., CLIP-style contrastive agreement) biases models toward modality alignment rather than incongruity detection. We present SCARF, a contradiction-aware framework that equips large multimodal models with explicit sarcasm cues and context-sensitive retrieval. SCARF constructs coarse scene cues and fine localized evidence via tag-constrained QA, then distills them with visual tokens into a [FUSION] control vector for the LLM; a label-contrastive retriever supplies type- and context-matched exemplars, and a local multi-view encoder surfaces micro-cues. With the same backbone and training data, SCARF attains 87.92% Acc/86.67% F1 on MMSD2.0 and 77.14% Acc/76.44% F1 zero-shot on XDMSD, outperforming a comparably fine-tuned LLaVA-1.5. Ablations show sarcasm clue fusion is the main driver of gains, and tag-constrained QA improves rationale grounding and reduces hallucinations. Yushan Pan, Ding Wang 0006, Wei Wang 0042, Xiaowei Huang 0001, Zhijie Xu |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2026 | DARE: Enriching Physical Dataflow Awareness for Macro Placement OptimizationabstractPhysical dataflow, which defines the detailed connections among cells and macros, is a critical yet underexplored factor in automatic macro placement. It becomes increasingly important for enabling intelligent design automation to minimize manual intervention and reduce design iterations. Existing macro or mixed-size placers with dataflow awareness primarily focus on intrinsic relationships among macros, overlooking the crucial influence of standard cell clusters on macro placement. To address this, we propose DARE, which extracts hidden connections between macros and standard cells and incorporates a series of algorithms to enrich dataflow awareness, integrating them into placement constraints for improved macro placement. To further optimize placement results, we introduce two fine-tuning steps: (1) congestion optimization by taking macro area into consideration, and (2) flipping decisions to determine the optimal macro orientation based on the extracted dataflow information. By integrating enhanced dataflow awareness into placement constraints and applying these fine-tuning steps, the proposed approach achieves an average 7.9% improvement in half-perimeter wirelength (HPWL) across multiple widely used benchmark designs compared to a state-of-the-art dataflow-aware macro placer. Additionally, it significantly improves congestion, reducing overflow by an average of 82.5%, and achieves improvements of 36.97% in Worst Negative Slack (WNS) and 59.44% in Total Negative Slack (TNS). The approach also maintains efficient runtime throughout the entire placement, incurring less than a 1.5% runtime overhead. These results show that the proposed dataflow-driven methodology, combined with the fine-tuning steps, provides an effective foundation for macro placement within the OpenROAD flow and can be further extended to other design flows in the future to enhance placement quality. Xiaotian Zhao, Yichen Cai 0004, Yushan Pan, Xinfei Guo |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2026 | Modality balancing network for pedestrian detection based on cross-modal compensation fusion and multimodal feature alignment
Zuhe Li, Ruochong Fu, Zhiyang Zhao, Penghao Ouyang, Zhijie Xu, Yushan Pan |
Vis. Comput. | 9 |
| 2025 | NAT-3D: A Non-Rigid Approach for Accurate 3D Tracking of Basic Surgical Instrumentsabstract3D tracking technology serves as an essential enabler in numerous domains, most notably within digital healthcare, where it supports critical procedures such as surgical navigation, trajectory planning, and high-fidelity surgical simulation, etc. However, due to the slender, textureless, and reflective characteristics of basic surgical instruments, coupled with the complexity of their diverse motion patterns, existing systems and methods face significant limitations in achieving accurate 3D tracking of these instruments. We propose NAT-3D, an effective and efficient non-rigid 3D tracking approach based on multi-modal region estimation, integrating kinematic structural constraints and nonrigid constraints. This method enables accurate tracking and dynamic mapping of a variety of basic surgical instruments, including scalpels, scissors, clamps, and forceps, without the need for additional markers. The tracking covers various movement modes, including rigid motion, local mechanical motion, and non-rigid composite motion. Extensive experiments demonstrate that our method outperforms previous algorithms in terms of robustness, accuracy, applicability, and real-time performance. In addition, we introduce a novel dataset for Basic Surgical Instrument Tracking (BSIT), which can serve as a benchmark for future related research. Zihan Deng, Jier Zhang, Yushan Pan, JunJun Pan, Zhijie Xu |
BIBM | 4 |
| 2025 | RDSA: A Robust Deep Graph Clustering Framework via Dual Soft Assignment
Yang Xiang 0009, Li Fan 0009, Tulika Saha, Xiaoying Pang, Yushan Pan, Haiyang Zhang 0004, Chengtao Ji |
DASFAA (3) | 5 |
| 2025 | The Impact of Artificial Intelligence Involvement in Video CreationabstractThe advancement of artificial intelligence (AI) has significantly transformed video creation, reshaping creators’ experiences, creative processes, and ethical considerations. This study employs Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA) to validate key dimensions, focusing on creators’ satisfaction, industry implications, and societal concerns related to AI involvement in video creation. Results from correlation analysis, ANOVA, T-tests, LSD post hoc tests, etc. reveal that creators with different characteristics hold varying perspectives on AI’s role in video production. While AI enhances creative efficiency and content quality, it also raises ethical and regulatory concerns. These findings provide valuable insights on how to balance AI’s benefits with ethical considerations, ensuring its sustainable and responsible application in the video creation industry. Haoran Qin, Yushan Pan, Xinda Li 0009, Zemeng Gu, Sanjeev Thaiyal |
INDIN | 2 |
| 2025 | Orion: A Multi-Agent Framework for Optimizing RAG Systems through Specialized Agent CollaborationabstractRetrieval-Augmented Generation (RAG) systems in enterprise environments face challenges including semantic overlap, ambiguous requirements, and precise knowledge retrieval difficulties.Despite Agent-based methodological advances, limitations persist in collaborative efficiency and contextual adaptability.To address these challenges,this paper introduces the Orion framework, a multi-Agent collaborative RAG architecture for Enterprise Guidance Q&A Systems.The framework integrates three specialized Agents: Information Amplification Specialist enhances content presentation to resolve similarity-based errors; Interaction Analyst optimizes query formulation to address non-expert articulation deficiencies; and Query Complexity Evaluator selects appropriate language models based on query characteristics.Empirical evaluation demonstrates 93.43% question-answering accuracy in enterprise scenarios, significantly outperforming existing systems.This multi-Agent framework enhances knowledge quality, improves retrieval precision, and generates responses better aligned with user requirements, establishing novel pathways for enterprise knowledge management and guidance question-answering systems. Xianxing Fang, Liangru Xie, Weibin Yang, Ruitao Zhang, Hao Wang 0003, Di Wu 0035, Yushan Pan |
Internetware | 8 |
| 2025 | Text-dominant multimodal perception network for sentiment analysis based on cross-modal semantic enhancements
Zuhe Li, Panbo Liu, Yushan Pan, Jun Yu 0011, Haoran Chen 0004, Hao Wang 0076 |
Appl. Intell. | 3 |
| 2025 | Sarcasm-GPT: advancing sarcasm detection with large language modelsabstractAbstract Sarcasm detection is a nuanced challenge in natural language processing, requiring deep understanding of textual and contextual cues. We present Sarcasm-GPT, a large language model-based model that integrates four key components: prompt template generation, retrieval-augmented generation, chain-of-thought generation, and a context fusion module. Together, these modules enrich contextual modeling, enable systematic reasoning, and seamlessly incorporate multimodal information, significantly boosting sarcasm detection accuracy. Extensive experiments, including ablation studies, confirm the complementary contributions of each module and demonstrate substantial performance gains over baselines. Additionally, our findings highlight the importance of context length in balancing interpretive accuracy and computational efficiency. The code is available at Sarcasm-GPT. Zuhe Li, Xiaojiang He, Yushan Pan |
Comput. J. | 5 |
| 2025 | MuST-GAN MFAS: Multi-semantic spoof tracer GAN with transformer layers for multi-modal face anti-spoofingabstractAbstract In the field of multi-modal face anti-spoofing (MFAS), where RGB, depth, and infrared data are integrated, remarkable advancements have been seen. However, despite the advancement, there still exist challenges when it comes to adaptability, particularly in dealing with unseen attacks. In this paper, a novel model called MuST-GAN MFAS is presented. This model employs a generative network that incorporates modality-specific encoders and transformer layers. It is significant that the model efficiently disentangles multi-semantic spoof traces by utilizing the power of cross-modal attention mechanisms and a transformer-based spoof trace generator. The training process involves bidirectional adversarial learning, ensuring identity consistency, intensity, center, and classification losses are taken into consideration. Through precise evaluations, it has been shown that the proposed model surpasses existing frameworks, showing remarkable performance when evaluating several modal samples. In the end, MuST-GAN MFAS makes an impressive contribution to the field of face anti-spoofing by offering results that are easy to interpret and emphasizing how important it is to learn multi-semantic spoof traces in order to improve generalization and adaptability to unseen attacks. The code is available at https://github.com/ZainUlAbideenMalik/Must-GAN-MFAS. Shu Liu 0002, Zain Ul Abideen 0009, Tongming Wan, Inzamam Shahzad, Waseem Abbas 0004, Yushan Pan |
Comput. J. | 6 |
| 2025 | From Routes to Ratings: Challenges and Strategies in Food Delivery Work
Haosu Xu, Yushan Pan, Shaoyu Cai |
Comput. Support. Cooperative Work. | 5 |
| 2025 | Text-guided multi-level interaction and multi-scale spatial-memory fusion for multimodal sentiment analysis
Xiaojiang He, Yanjie Fang, Zuhe Li, Chenguang Yang 0001, Hao Wang 0003, Yushan Pan |
Neurocomputing | 8 |
| 2025 | Multimodal sentiment analysis based on disentangled representation learning and cross-modal-context association mining
Zuhe Li, Panbo Liu, Yushan Pan, Weiping Ding 0001, Jun Yu 0011, Haoran Chen 0004, Hao Wang 0003 |
Neurocomputing | 3 |
| 2025 | Empowering couriers: balancing algorithmic control and autonomy for enhanced well-beingabstractAbstract This paper explores the complex socio-technical dynamics that shape the work experience of food delivery workers in China through a qualitative study of 28 couriers. Our findings reveal a serious disconnect between algorithmic systems and workers’ lived experiences. We identify five key tensions in this algorithmic labour environment: (1) complementarity and conflict between algorithmic logic and workers’ experiential knowledge; (2) business models that shape technological systems without adequately accounting for task complexity; (3) rating systems that simultaneously quantify performance but creates ”service inflation” and psychological stress; (4) privacy concerns create a tension between service efficiency and information protection; and (5) complaint handling mechanisms systematically shift risks to workers through disproportionate penalties and defense costs. These findings contribute to the human–computer interaction literature by showing how algorithmic management embodies power structures rather than being a neutral tool, how workers’ adaptability constitutes valuable but invisible labor, how the trade-off between privacy and efficiency manifests differently in the Chinese cultural context, and how procedural justice issues extend beyond technical design. Finally, we suggest the implications of designing more equitable sociotechnical systems, namely better integrating workers’ tacit knowledge, matching business models to technical capabilities, developing culturally appropriate privacy solutions, and establishing more balanced evaluation and grievance resolution processes. Haosu Xu, Zhijie Xu, Yushan Pan |
Interact. Comput. | 5 |
| 2025 | Back to fundamentals: Low-level visual features guided progressive token pruning
Yizhuo Liang 0002, Qingpeng Li, Xinfei Guo, Di Wu 0035, Hao Wang 0003, Yushan Pan |
J. Syst. Archit. | 8 |
| 2025 | Dual-stream autoencoder for channel-level multi-scale feature extraction in hyperspectral unmixing
Yuquan Gan, Yushan Pan |
Knowl. Based Syst. | 6 |
| 2025 | Representation distribution matching and dynamic routing interaction for multimodal sentiment analysis
Zuhe Li, Zhenwei Huang, Xiaojiang He, Jun Yu 0011, Haoran Chen 0004, Chenguang Yang 0001, Yushan Pan |
Knowl. Based Syst. | 7 |
| 2025 | Text-guided deep correlation mining and self-learning feature fusion framework for multimodal sentiment analysis
Xiaojiang He, Baojie Qiao, Zuhe Li, Yushan Pan |
Knowl. Based Syst. | 6 |
| 2025 | Graph-Aware Hybrid Encoding for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification faces critical challenges in effectively modeling the intricate spectral-spatial structures and non-Euclidean relationships. Traditional methods often struggle to simultaneously capture local details, global contextual dependencies, and graph-structured correlations, leading to limited classification accuracy. To address the above issues, this letter proposes a Graph-Aware Hybrid Encoding (GAHE) framework. To fully exploit the spectral-spatial characteristics and graph structural dependencies inherent in HSI, the proposed method is structured into three key components: a multi-scale selective graph-aware attention module, a hybrid projection encoding module, and a graph sensitive aggregation module. The three modules work in a complementary manner to progressively refine and enhance feature representations across multiple scales and modalities. Comparing with advanced classification methods,the experimental results demonstrate that the proposed GAHE method shows better classification performance. Yuquan Gan, Zhijie Xu, Yushan Pan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Efficient 3D human pose estimation via spatio-temporal graph transformer with token pruning
Zuhe Li, Hongyang Chen 0001, Fengqin Wang, Qidong Liu 0007, Yushan Pan |
Multim. Syst. | 6 |
| 2025 | Scale-Selectable Global Information and Discrepancy Learning Network for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis and depression detection are pivotal for advancing human-computer interaction, yet significant challenges remain. First, the limited extraction of global contextual information within individual modalities risks the loss of modal-specific features. Second, existing methods often prioritize unaligned textual interactions, neglecting critical inter-modal discrepancies. To address these issues, we propose the Scale-Selectable Global and Discrepancy Learning Network (SSGDL), an innovative framework that integrates two core modules: the Cross-Shaped Dynamic Scale Attention Module (CSDSA) and the Primary-Secondary modal Discrepancy Learning Module (PS-MDL). The CS-DSA dynamically selects scales and employs cross-shaped attention to capture comprehensive global context and intricate internal correlations, effectively producing a fused modal representation. Meanwhile, the PS-MDL designates the fused modal as primary and utilizes cross-attention mechanisms to learn discrepancy representations between it and other modalities (textual, acoustic, and visual). By leveraging intermodal discrepancies, SSGDL achieves a more nuanced and holistic understanding of emotional content. Extensive experiments on three benchmark multimodal sentiment analysis datasets (MOSI, MOSEI, SIMS) and a depression detection dataset (AVEC2019) demonstrate that SSGDL consistently outperforms state-of-theart approaches, setting a new benchmark for multimodal affective computing. Xiaojiang He, Yushan Pan, Xinfei Guo, Zhijie Xu, Chenguang Yang 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | DSANet: Dynamic and Structure-Aware GCN for Sparse and Incomplete Point Cloud LearningabstractLearning 3-D structures from incomplete point clouds with extreme sparsity and random distributions is a challenge since it is difficult to infer topological connectivity and structural details from fragmentary representations. Missing large portions of informative structures further aggravates this problem. To overcome this, a novel graph convolutional network (GCN) called dynamic and structure-aware NETwork (DSANet) is presented in this article. This framework is formulated based on a pyramidic auto-encoder (AE) architecture to address accurate structure reconstruction on the sparse and incomplete point clouds. A PointNet-like neural network is applied as the encoder to efficiently aggregate the global representations of coarse point clouds. On the decoder side, we design a dynamic graph learning module with a structure-aware attention (SAA) to take advantage of the topology relationships maintained in the dynamic latent graph. Relying on gradually unfolding the extracted representation into a sequence of graphs, DSANet is able to reconstruct complicated point clouds with rich and descriptive details. To associate analogous structure awareness with semantic estimation, we further propose a mechanism, called structure similarity assessment (SSA). This method allows our model to surmise semantic homogeneity in an unsupervised manner. Finally, we optimize the proposed model by minimizing a new distortion-aware objective end-to-end. Extensive qualitative and quantitative experiments demonstrate the impressive performance of our model in reconstructing unbroken 3-D shapes from deficient point clouds and preserving semantic relationships among different regional structures. Yushi Li, George Baciu, Rong Chen 0003, Chenhui Li 0001, Hao Wang 0003, Yushan Pan, Weiping Ding 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | "Please Be Nice": Robot Responses to User Bullying - Measuring Performance Across Aggression LevelsabstractAs robots become integral to public services, addressing harmful user behaviors like bullying is crucial. Existing research often overlooks the gradual nature of human bullying. This study fills this gap by exploring how robots can counter bullying through optimized responses. Using a simulated human-robot interaction study, we manipulated robot response behaviors and styles across escalating bullying severity. Results show that empathetic verbal responses promptly reduce users’ bullying tendencies by eliciting remorse and redirecting attention to social awareness. However, users’ underlying dispositions may override these reflexive reactions, emphasizing the need for a holistic understanding. In conclusion, a comprehensive approach is essential, involving immediate reaction optimization, emotional state assessment, and ongoing behavioral adjustment through empathetic dialogue. By implementing such strategies, we can transform human-robot relationships from potential bullying situations to harmonious interactions. This study provides an empirical foundation for response protocols that discourage bullying and enhance mutual understanding. Di Wu 0035, Hao Wang 0003, Yushan Pan |
CHI | 5 |
| 2024 | Document Set Expansion with Positive-Unlabeled Learning Using Intractable Density EstimationabstractThe Document Set Expansion (DSE) task involves identifying relevant documents from large collections based on a limited set of example documents. Previous research has highlighted Positive and Unlabeled (PU) learning as a promising approach for this task. However, most PU methods rely on the unrealistic assumption of knowing the class prior for positive samples in the collection. To address this limitation, this paper introduces a novel PU learning framework that utilizes intractable density estimation models. Experiments conducted on PubMed and Covid datasets in a transductive setting showcase the effectiveness of the proposed method for DSE. Code is available from https://github.com/Beautifuldog01/Document-set-expansion-puDE. Haiyang Zhang 0004, Qiuyi Chen, Yanjie Zou, Jia Wang 0009, Yushan Pan, Mark Stevenson 0001 |
LREC/COLING | 5 |
| 2024 | Language-based Audio Retrieval with GPT-Augmented Captions and Self-Attended Audio ClipsabstractWith the explosion of user-generated content in recent years, efficient methods for organizing multimedia databases based on content and retrieving relevant items have become essential. Language-based audio retrieval seeks to find relevant audio clips based on natural language queries. However, there exists a scarcity of datasets specifically developed for this task. Moreover, the language annotations often carry biases, leading to unsatisfactory retrieval accuracy. In this work, we propose a novel framework for language-based audio retrieval that aims to: 1) utilize GPT-generated text to augment audio captions, thereby improving language diversity; 2) employ audio self-attention mechanisms to capture intricate acoustic features and temporal dependencies. Experiments conducted on two public datasets, containing both short- and long-term audios, demonstrate that our framework can achieve significant performance improvements compared with other methods. Specifically, the proposed framework can achieve a 27% increase in mean average precision (mAP) on the Clotho dataset, and a 31% improvement in mAP on the AudioCaps dataset compared with the baseline. Fuyu Gu, Yiyan Xu, Yushan Pan, Shengchen Li, Haiyang Zhang 0004 |
CSCWD | 5 |
| 2024 | A novel model for product defect detection based on the automatic aspect term extractionabstractIndustry 4.0 is a global modernization movement in the manufacturing industry, heralding a new paradigm of digitalization, decentralized control, and autonomy in the realm of the manufacturing industry. From the perspective of quality management, Product Defect Detection (PDD) is an important industrial application throughout the entire product life-cycle management, which aims to detect the defects of products. Product defect, which essentially involves safety-related defect and performance defect, has been drawing widespread attention from business and academia due to their relevance to failure costs and the user experience. With the rapid development of social media, performance defects are increasingly discussed in online reviews and forums, and there also exist some text-mining-based methods to extract the defect information. However, the existing methods suffer from the limitation of expert dependence and failure detection of unseen defects. To fill these gaps, we first creatively introduce the semi-structured eWOM (electronic word-of-mouthh) data to avoid manual work. Secondly, we designed the automatic aspect term extraction (ATE) module to extract the product components to tackle the problem of the predefined corpus. Besides, detailed hierarchical defect information can also be provided to support the manufacturer in locating the defect component. Thirdly, an aspect-level text classification model is introduced with the advantage of the local-context attention mechanism. A case study in the automotive industry is provided to validate the effectiveness of the proposed model, and a wide range of experiments are offered to illustrate the high performance of the proposed model. Zhongyun Li, Zongyang Liu, Yushan Pan |
CSCWD | 5 |
| 2024 | Elastic EDA: Auto-Scaling Cloud Resources for EDA Tasks via Learning-based ApproachesabstractUtilizing cloud EDA for chip design allows access to on-demand high-performance computing (HPC) resources, significantly reducing development time and costs by eliminating the need for costly on-site infrastructure. Despite its benefits, cloud EDA faces significant challenges, primarily the lack of an effective cost model. A key issue is the absence of a mechanism for designers to accurately gauge the characteristics of their EDA jobs in cloud environment, as the design process involves a multitude of EDA tools and steps, often leading to the over or underestimation of needed computational resources. This problem is exacerbated by the varying computational demands of different designs and constraints. To bridge this knowledge gap, we introduce Elastic EDA, a methodology that harnesses machine learning (ML) to understand the characteristics of a design and its early stages, and to predict the computational needs for subsequent phases throughout the entire EDA flow. This approach effectively aligns design behaviors with computational resources, providing cost-efficient solutions for various cloud EDA scenarios. Compared to previous ML-based predictive frameworks for cloud EDA, the proposed method achieves over 60% higher prediction accuracy and supports various elastic computing environments, maximizing the efficiency of cloud re-sources. Compared to various baseline scheduling configurations in the cloud environment, the proposed framework achieves over 16% mean runtime improvement. Linyu Zhu, Shaogang Hao, Yushan Pan, Xinfei Guo |
ICCD | 4 |
| 2024 | Joint Multiscale Spatial-Frequency Domain Network for Oriented Object Detection in Remote Sensing ImagesabstractThe detection of oriented object detection in remote sensing images remains a daunting challenge due to their complex backgrounds, various sizes, and especially arbitrary orientations. However, most existing methods only model the structural features of images in the spatial domain, while the horizontal convolution kernels limit the model’s ability to perceive object direction information. Furthermore, frequency features contain rich information about scale, texture and angle, which can be a good complement to the spatial features. Inspired by this, we propose a multiscale spatial-frequency domain network (MSFN) to utilize spatial-frequency information for oriented object detection, which can be integrated into any CNN architectures seamlessly and perform end-toend training easily. Besides, we design a channel alignment feature fusion module (CA-FFM) to address the problem of fusing low-level texture and high-level semantic features with a significant difference in channel dimension. Experimental results on HRSC2016 and SSDD dataset demonstrate the effectiveness of the proposed method. Yushan Pan, Yang Xu 0006, Zebin Wu 0001, Zhihui Wei |
IGARSS | 1 |
| 2024 | LB-KBQA: Large-language-model and BERT based Knowledge-Based Question and Answering SystemabstractGenerative Artificial Intelligence (AI), because of its emergent abilities, has empowered various fields, one typical of which is large language models (LLMs). One of the typical application fields of Generative AI is large language models (LLMs), and the natural language understanding capability of LLM is dramatically improved when compared with conventional AI-based methods. The natural language understanding capability has always been a barrier to the intent recognition performance of the Knowledge-Based-Question-and-Answer (KBQA) system, which arises from linguistic diversity and the newly appeared intent. Conventional AI-based methods for intent recognition can be divided into semantic parsing-based and model-based approaches. However, both of the methods suffer from limited resources in intent recognition. To address this issue, we propose a novel KBQA system based on a Large Language Model(LLM) and BERT (LB-KBQA). With the help of generative AI, our proposed method could detect newly appeared intent and acquire new knowledge. In experiments on financial domain question answering, our model has demonstrated superior effectiveness. Zhongyun Li, Yushan Pan, Zhiman Zhang |
INDIN | 3 |
| 2024 | Leveraging Large Language Models for QA Dialogue Dataset Construction and Analysis in Public Services
Chaomin Wu, Di Wu 0035, Yushan Pan, Hao Wang 0003 |
NLPCC (1) | 3 |
| 2024 | WalkFormer: Point Cloud Completion via Guided WalksabstractPoint clouds are often sparse and incomplete in real-world scenarios. The prevailing methods for point cloud completion typically rely on encoding the partial points and then decoding complete points from a global feature vector, which might lose the existing patterns and elaborate structures. To address these issues, we propose WalkFormer, a novel approach to predict complete point clouds through a partial deformation process. Concretely, our method samples locally dominant points based on feature similarity and moves the points to form the missing part. Since these points maintain representative information of the surrounding structures, they are appropriately selected as the starting points for multiple guided walks. Furthermore, we design a Route Transformer module to exploit and aggregate the walk information with topological relations. These guided walks facilitate the learning of long-range dependencies for predicting shape deformation. Qualitative and quantitative evaluations demonstrate that our proposed approach achieves superior performance compared to state-of-the-art methods in the 3D point cloud completion task. Mohang Zhang, Yushi Li, Rong Chen 0003, Yushan Pan, Rong Xiang |
WACV | 4 |
| 2024 | C²DR: Robust Cross-Domain Recommendation based on Causal DisentanglementabstractCross-domain recommendation aims to leverage heterogeneous information to transfers knowledge from a data-sufficient domain (source domain) to a data-scarce domain (target domain). Existing approaches mainly focus on learning single-domain user preferences and then employ a transferring module to obtain cross-domain user preferences, but ignore the modeling of users' domain specific preferences on items. We argue that incorporating domain-specific preferences from the source domain will introduce irrelevant information that fails to the target domain. Additionally, directly combining domain-shared and domain-specific information may hinder the target domain's performance. To this end, we propose C^2DR, a novel approach that disentangles domain-shared and domain-specific preferences from a causal perspective. Specifically, we formulate a causal graph to capture the critical causal relationships based on the underlying recommendation process, explicitly identifying domain-shared and domain-specific information as causal irrelevant variables. Then, we introduce disentanglement regularization terms to learn distinct representations of the causal variables that obey the independence constraints in the causal graph. Remarkably, our proposed method enables effective intervention and transfer of domain-shared information, thereby improving the robustness of the recommendation model. We evaluate the efficacy of C^2DR through extensive experiments on three real-world datasets, demonstrating significant improvements over state-of-the-art baselines. Menglin Kong, Jia Wang 0009, Yushan Pan, Haiyang Zhang 0004, Muzhou Hou |
WSDM | 3 |
| 2024 | Anomaly Metrics on Class Variations For Face Anti-SpoofingabstractAbstract In face anti-spoofing tasks, distinguishing between live and spoof faces across different data domains presents challenges due to inter-class similarities, intra-class variations and unknown spoof patterns. This hampers generalization in real-world applications. To address this, we propose a novel convolutional neural network framework that utilizes spatial-frequency cues for 2D and 3D attacks. Furthermore, we introduce compact anomaly metrics and design three anomaly metrics-based supervisions from the perspective of Reed-Xiaoli anomaly detection, aiming to tackle the challenge posed by unknown attacks. Thanks to our proposed spatial frequency factorization network and its frequency-related supervisions, the spoofing cues are significantly enhanced, resulting in remarkable improvements in our experimental results. These outcomes demonstrate that our proposed framework achieves state-of-the-art performance on both monocular and multi-spectral benchmark datasets. Bing Gong, Kai Che, Jieming Ma, Yushan Pan |
Comput. J. | 5 |
| 2024 | Hierarchical denoising representation disentanglement and dual-channel cross-modal-context interaction for multimodal sentiment analysis
Zuhe Li, Zhenwei Huang, Yushan Pan, Jun Yu 0011, Haoran Chen 0004, Di Wu 0035, Hao Wang 0003 |
Expert Syst. Appl. | 3 |
| 2024 | Feasibility and performance enhancement of collaborative control of unmanned ground vehicles via virtual reality
Ziming Li 0003, Jialin Wang 0002, Yushan Pan, Lingyun Yu 0001, Hai-Ning Liang |
Pers. Ubiquitous Comput. | 4 |
| 2024 | Correction to: Feasibility and performance enhancement of collaborative control of unmanned ground vehicles via virtual reality
Ziming Li 0003, Jialin Wang 0002, Yushan Pan, Lingyun Yu 0001, Hai-Ning Liang |
Pers. Ubiquitous Comput. | 4 |
| 2024 | A Mask Guided Oriented Object Detector Based on Rotated Size-Adaptive Tricube KernelabstractOriented object detection is an important research topic in remote sensing. The detection of oriented objects in remote sensing images remains a daunting challenge due to their complex backgrounds, various sizes, diverse aspect ratios, and especially arbitrary orientations. In recent years, keypoint-based anchor-free object detectors have demonstrated outstanding performance in this field. However, in current anchor-free detectors, object keypoints are primarily generated using the Gaussian kernel function, which assumes a circular form. This representation falls short in accurately conveying an object’s size and orientation. To address the aforementioned issue, this paper proposes a keypoint-based oriented object detector called MRSDet, which innovatively adopts the Tricube kernel, scales and rotates it, to better generate the center keypoint heatmap of the object. Besides, to improve the model’s detection performance on oriented objects and improve its ability to perceive object keypoints and boundary boxes, we also design a large receptive field mask module (LRFM), which is based on large convolution kernel decomposition and semantic segmentation masks. Taking the BBAVectors method as a baseline, we conduct experiments on multiple types of remote sensing datasets such as HRSC2016, UCAS-AOD and SSDD+ to verify the effectiveness and generalizability of the proposed method. Yushan Pan, Yang Xu 0006, Zebin Wu 0001, Zhihui Wei, Javier Plaza, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Channel Self-Attention Based Multiscale Spatial-Frequency Domain Network for Oriented Object Detection in Remote Sensing ImageryabstractThe detection of oriented objects in remote sensing images remains a daunting challenge due to their complex backgrounds, various sizes, and especially arbitrary orientations. However, most of the existing methods only model the structural features of the images in the spatial domain, while the horizontal convolution kernels limit the model’s ability to perceive object direction information. Furthermore, the frequency features contain rich information about scale, texture, and angle, which can be a good complement to the spatial features. Inspired by this, we propose a multiscale spatial-frequency domain network (MSFN) to utilize spatial-frequency information for oriented object detection, which can be integrated into any convolutional neural network (CNN) architectures seamlessly and perform end-to-end training easily. Firstly, multiscale Haar wavelet transforms are leveraged to extract the multiscale frequency domain features from the image. Subsequently, channel alignment feature fusion module (CA-FFM) is proposed to fuse the high-level semantic features extracted by CNN with the low-level texture features extracted by the wavelet transform in multiscale. Finally, a channel self-attention (CSA)-based spatial-frequency feature perception module (SFPM) is designed to perform self-attention weighted aggregation on the fused features along the channel dimension, thereby constructing a novel spatial-frequency feature extraction backbone network for oriented object detector in remote sensing images. Experimental results on the DOTA and HRSC2016 datasets validate the effectiveness and universality of the proposed method. Yang Xu 0006, Yushan Pan, Zebin Wu 0001, Zhihui Wei, Tianming Zhan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | A Study of Zooming, Interactive Lenses and Overview+Detail Techniques in Collaborative Map-based TasksabstractThe support for multi-focus data exploration is vital in collaborative visualization. In these scenarios, which often involve multiple devices and large displays, users may focus on specific information on their individual screens while also sharing contextual views with others. While many visualization techniques developed for single-user applications can be adapted for use in collaborative settings, little research has been done on how to design adaptive versions of these techniques or how they may impact collaborative tasks involving large datasets. In this work, we perform a comparative study of three collaborative visualization techniques (Zooming, Interactive lenses and Overview+Detail) on large displays in three map-based visualization tasks (Exploration, Comparison and Spatial Memorizing). These three collaborative techniques draw on three different classical visualization techniques in a single-user setting. Our results show that these techniques have different impacts on users’ task performance and preferences. The collaborative Overview+Detail technique benefits users most in supporting Spatial Memorizing. Closely coupled groups prefer collaborative Zooming in Target Exploration. Based on these results, we further discuss the design of collaborative visualization techniques and propose suggestions for adapting classical single-user visualization techniques to a collaborative setting. Yu Liu 0077, Yushan Pan, Yue Li 0023, Hai-Ning Liang, Paul Craig, Lingyun Yu 0001 |
PacificVis | 3 |
| 2021 | Reflexivity of Account, Professional Vision, and Computer-Supported Cooperative Work: Working in the Maritime DomainabstractComputer-supported cooperative work (CSCW) researchers have plenty to say about designing through texts. However, the relationships between implementation and design are often misunderstood by engineering project systems developers who design collaborative systems, among others. This problem is not new; how to utilize CSCW insights effectively and correctly in engineering projects has long been a concern in the CSCW community. By reviewing year-long, multiple-site ethnographic studies in the maritime domain conducted since 2015, this paper reflects on the "reflexivity of account" and "professional vision" as complementary concepts for ensuring that actors' actions and reasonings are designed in such a way that actors-in this case, maritime operators, systems developers, educators, policymakers, and shipowners-are accountable in and through their member groups. Rather than generalizing my findings and my role in the maritime domain as explicit knowledge to help other CSCW researchers in studying engineering projects, the goal of this paper is to seek intersubjective knowledge of my fieldwork to trigger the utilization of CSCW insights as a fundamental basis for facilitating systems design. The aim is to shape and reshape systems in line with what those actors do in reality. Yushan Pan |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2016 | Cooperative systems for marine operations using actor-network design: A discussionabstractAfter discussing the roles of marine engineers and interaction designers from computer science (CS)/information systems (IS), I investigate issues with the aim of promoting better cooperative systems for maritime industries. As part of my research project, I have observed marine operations and interviewed bridge operators, captains, and crewmembers, as well as engine engineers in their workplace, in order to investigate how cooperative work occurs in current marine operational systems. Against the background of computer-supported cooperative work (CSCW), I look into how to improve systems to support safety-critical operations for offshore services at sea. Most safety-critical operations occur as part of cooperative operations and interactions between operators and systems. Engineers and interaction designers should cooperate to design cooperative systems with users in the early stages. However, there is not a common understanding between engineers and interaction designers regarding the meaning of cooperative work. In order to improve computer systems to support marine cooperative work, a new understanding of marine technology is necessary where design needs are met through the integration of engineering design, CS/IS theoretical thinking, and work practices. Human operators, systems, and materials used in workspaces are actors deeply lived in work environments; hence, I discuss a means of adopting Bannon's concept of human actors toward the development of an actor-network design of cooperative systems. Yushan Pan |
CSCWD | 1 |
| 2013 | Tactile Cues For Ship Bridge OperationsabstractCurrent modes of conveying operational information on ship bridges are mainly in the form of visual and auditory sensory inputs. During safety critical situations, such as dynamic positioning (DP) operation, reliance on these two senses may be insufficient. In a DP operation the role of the DP operator is critical and most incidents happen due to lack of operator s situational awareness. Therefore exploitation of other sensory inputs in addition to visual and auditory must be investigated. This paper surveys recent research on response times of vibro-tactile, visual and auditory cues. The survey concludes that tactile cue always has shorter response time compared to other stimuli. And its combination with other cues such as visual and auditory can enhance the effectiveness of response time. Therefore new ship bridge developers should take this knowledge into account to increase design efficiency. Yushan Pan, Sathiya Kumar Renganayagalu, Sashidharan Komandur |
ECMS | 1 |
| 2012 | A comparative review of i∗-based and use case-based security modelling initiativesabstractSecurity requirements elicitation and modelling are integral for the successful development of secure systems. However, there are a lot of similar yet not identical approaches that currently exist for security requirements modelling, which is confusing for researchers and practitioners hence some characterisation will be useful to give a better overview and understanding of advantages and disadvantages of various approaches. This paper provides a comparative review of i*-based and use case - based security modelling initiatives, using a characterisation framework with several dimensions. Our findings show that both categories of initiatives have significant conceptual similarities in the aspect of modelling language and method process, and coverage of security requirements modelling notions. They have conceptual differences in the aspect of: representation perspective, kind of security requirements engineering activities that are supported, the quality of specification that is generated and the specification techniques used, and the degree of support for software evolution. Olawande J. Daramola, Yushan Pan, Péter Kárpáti, Guttorm Sindre |
RCIS | 2 |