Haiyan Wu

dblp:11/3756 · DBLP profile ↗
← Back
48ranked-venue papers
18as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 12 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 10 since 2021Systems, architecture and hardware · 7 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 The Silent Amplifier: In-Context Examples Fuel Bias in Large Language Models
abstract
In-context learning (ICL) has proven to be adept at adapting large language models (LLMs) to downstream tasks without parameter updates, based on a few demonstration examples. Prior work has found that the ICL performance is susceptible to the selection of examples in prompt and made efforts to stabilize it. However, existing example selection studies ignore the ethical risks behind the examples selected, such as gender and race bias. In this work, we conduct extensive experiments and discover that (1) example selection with high accuracy does not mean low bias; (2) example selection for ICL may amplify the biases of LLMs; (3) example selection contributes to spurious correlations of LLMs. Based on the above observations, we propose the Remind with Bias-aware Embedding (ReBE), which removes the spurious correlations through contrastive learning and obtains bias-aware embedding for LLMs based on prompt tuning. Finally, we demonstrate that ReBE effectively mitigates biases of LLMs without significantly compromising accuracy and is highly compatible with existing example selection methods.
Jiashi Gao, Junlei Zhou, Jiaxin Zhang 0007, Quanying Liu, Haiyan Wu, Xin Yao 0001, Xuetao Wei
AAAI6
2026 GeWu: A Culturally-Grounded Chinese Benchmark for Multi-Stage Social Bias Evaluation in Large Language Models
abstract
With the rapid deployment of Chinese large language models (LLMs), culturally-grounded bias evaluation remains understudied due to the dominance of English benchmarks and simplistic Chinese scenarios. To address this, we propose GeWu, a comprehensive benchmark featuring a culturally-aware dataset of 60,192 questions spanning 14 social groups with fine-grained Chinese contexts, significantly exceeding existing resources in breadth and depth. Our two-stage evaluation first quantifies bias via multiple-choice questions using a novel probability-based scoring mechanism to sensitively capture bias tendencies, distilling high-bias scenarios into GeWu-1K. This refined subset then enables multi-turn dialogue evaluations for in-depth analysis under realistic conditions. Experiments reveal that GeWu effectively exposes social biases in state-of-the-art Chinese LLMs, with 13.93% of scenarios eliciting universal bias across all models. This highlights persistent challenges and provides actionable insights for bias mitigation in Chinese contexts.
Jiashi Gao, Jiaxin Zhang 0007, Haiyan Wu, Xin Yao 0001, Xuetao Wei
AAAI6
2026 GPTM: Gaussian Prior-Guided Temporal Modulation for DiffusionDet-Based Tiny Object Detection in UAV Imagery
Haiyan Wu
ICIC (12)1
2026 Conflict-Aware RAG: Multi-Stage Learning with Conflict Signals for Robust Retrieval-Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) effectively mitigates hallucinations and knowledge gaps in Large Language Models (LLMs) for knowledge-intensive tasks by incorporating external web-based knowledge. However, when integrating diverse yet potentially conflicting web-sourced information, RAG systems are prone to knowledge conflicts that manifest as incorrect or inconsistent model behaviors, ultimately leading to unreliable responses. To address this challenge, this paper proposes Conflict-Aware RAG, a general training framework that leverages the model's inherent conflict-sensing capability to build a more robust RAG system via phased optimization. At the core of this framework lies ConScore, a conflict signal that quantifies the model's awareness of potential knowledge conflicts by comparing generative probabilities across distinct knowledge sources. This signal then guides both the construction of training data and a multi-stage optimization workflow: In the Supervised Fine-Tuning (SFT) stage, conflict features are employed to select representative distracting documents, laying the groundwork for core RAG capabilities; in the Direct Preference Optimization (DPO) stage, high-quality preference pairs are constructed using the conflict signal to boost the model's robustness against distracting knowledge; and in the Reranking stage, conflict confidence and information gain are integrated to synergistically optimize the collaboration mechanism between the retriever and LLM. Experiments on six knowledge-intensive question answering (QA) datasets demonstrate that Conflict-Aware RAG significantly outperforms mainstream baselines. Further ablation studies and quantitative analyses validate the method's stability and generalization, laying the foundation for robust RAG systems.
Haiyan Wu, Chaoqun Sun, Chengxiong Lu, Zhiqiang Zhang 0010
WWW1
2026 Multidimensional Contextual Knowledge Inference Model for Sarcasm Detection
abstract
Sarcasm detection contributes to the understanding of the contrast between the literal meaning of an utterance and the true intention of the speaker, and is considered to be part of the challenge in sentiment analysis. Most models usually focus on prompted inference, ignoring the knowledge hallucination (i.e., the generation of factually incorrect or fabricated information) and inference mistakes generated by large models, which affects the detection accuracy of sarcastic semantics. To solve these complex problems, we propose a novel multidimensional contextual knowledge inference (MCKI) model, which further enhances the contextual inference capability of the model by introducing the multihop chain of thought (CoT) inference technique, combining knowledge enhancement techniques and self-consistency mechanism to better capture the underlying intent of sarcastic text. Specifically, we first utilize multihop context inference techniques to excavate fine-grained contextual and emotional cues in satire through a stepwise inference process. Second, contextual knowledge augmentation is employed to motivate the model to capture deep semantics. Finally, a self-consistency mechanism is adopted to ensure that the generated inference paths are consistent across multiple perspectives, thus improving the accuracy and stability of sarcasm detection. Experimental results illustrate that compared to the state-of-the-art baseline, our model achieves significant improvements on the three baseline datasets, especially on the X dataset by 2.4%, which demonstrates the superior performance of our Multidimensional contextual knowledge Inference model for sarcastic detection.
Zhiqiang Zhang 0010, Bing Li 0027, Haiyan Wu, Yuankang Sun, Haimiao Mo
IEEE Trans. Comput. Soc. Syst.4
2026 TrustSyn: Augmenting LLMs With Constituency- Structured Dependency Knowledge for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) constitutes a critical subtask within affective computing, whose central challenge involves the accurate and efficient identification of sentiment polarity associated with specific aspect terms in review sentences. Although syntactic knowledge has demonstrated significant benefits in traditional ABSA models, existing approaches based on large language models (LLMs) have largely overlooked such structural information and often fail to comprehensively model both implicit and explicit sentiment expressions. To bridge this gap, we propose TrustSyn, a novel framework designed to enhance LLMs with trustworthy, constituency-structured dependency knowledge for ABSA. Specifically, the input sentences are first parsed using both dependency and constituency parsers. The resulting syntactic information is then restructured into a unified and reliable representation through a trustworthy syntax integration process. This structured knowledge is formalized and injected into LLMs to augment their comprehension of aspect sentiment associations. To the best of our knowledge, this is the first work to integrate constituency-informed dependency structures into LLMs for ABSA. Finally, experimental results demonstrate that TrustSyn consistently outperforms state-of-the art models across five benchmark datasets. Further ablation studies and analyses confirm its robustness and strong generalization capability.
Haiyan Wu, Chaoqun Sun, Chengxiong Lu, Jianyong Wang 0001, Zhiqiang Zhang 0010
IEEE Trans. Knowl. Data Eng.1
2025 LLMs Trust Humans More, That's a Problem! Unveiling and Mitigating the Authority Bias in Retrieval-Augmented Generation
abstract
Yuxuan Li, Xinwei Guo, Jiashi Gao, Guanhua Chen, Xiangyu Zhao, Jiaxin Zhang, Quanying Liu, Haiyan Wu, Xin Yao, Xuetao Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jiashi Gao, Guanhua Chen 0001, Xiangyu Zhao 0001, Jiaxin Zhang 0007, Quanying Liu, Haiyan Wu, Xin Yao 0001, Xuetao Wei
ACL (1)8
2025 Generative Diffusion Model-Enhanced Federated Fine-Tuning for Resource-Aware Edge Intelligence
abstract
Edge devices increasingly require efficient, on-device intelligence for diverse applications in IoT networks. In order to bring the advanced capabilities of large foundation models directly to the point of data generation, there is a growing interest in deploying these models on edge devices. However, due to their inherent resource constraints and the diverse, heterogeneous nature of the data and tasks they encounter, deploying large foundation models directly on these devices remains a significant challenge. To address these challenges, we propose a novel Federated Learning Fine-Tuning (FLFT) framework that leverages adapter-based fine-tuning with a similarity-driven selection mechanism, enabling personalized model adaptation with minimal computational overhead. Furthermore, we introduce the Diffusion-based Soft Actor-Critic (FTFL2DSAC) algorithm, which optimizes real-time resource allocation by balancing energy consumption and latency across heterogeneous edge devices. Our experiments on CIFAR-100 using a pre-trained multimodal model demonstrate that FLFT achieves 82.5% accuracy while reducing model parameters by 14%, outperforming baseline methods with faster convergence and enhanced stability in complex environments.
Haiyan Wu, Wenji He, Lin Du 0006, Xiaoxu Ren, Tianhao Ouyang, Haipeng Yao
IWCMC1
2025 Mitigating Stereotypes in Text-to-Image Generation: A Novel Perspective of Selective Neural Suppression
abstract
Text-to-Image (T2I) diffusion models exhibit concerning tendencies to generate harmful imagery that perpetuates social biases and stereotypes, posing significant ethical risks in real-world applications. While existing mitigation approaches predominantly employ black-box methodologies through dataset augmentation or constrained fine-tuning, they face critical limitations, including high data acquisition costs and potential exacerbation of stereotypes during model retraining. Inspired by neuroscience principles where neurological dysfunction often stems from aberrant neural activation patterns, we propose a novel framework, StereoClinic, targeting the root cause of stereotype generation through direct neural intervention. Our solution introduces two synergistic components: Diffusion Deep Taylor Decomposition (DDTD) for precisely localizing stereotype-related neurons via Layer-wise Relevance Propagation (LRP) attribution analysis, and Stereotype Neuron Suppression (SNS) implementing targeted activation damping to neutralize bias propagation. Through extensive empirical evaluations across multiple bias dimensions, we demonstrate that our method achieves significant stereotype mitigation without compromising image quality or requiring additional training data. This neuro-inspired approach establishes a new paradigm for model interpretability and ethical alignment in generative AI systems.
Junlei Zhou, Jiashi Gao, Haiyan Wu, Quanying Liu, Xiangyu Zhao 0001, Hongxin Wei, Xin Yao 0001, Xuetao Wei
ACM Multimedia4
2025 High-quality parameterization of single-axis swept volumes for isogeometric analysis
Jinlan Xu, Jiakai Yu, Haiyan Wu
Comput. Graph.3
2024 Feature-preserving quadrilateral mesh Boolean operation with cross-field guided layout blending
Haiyan Wu, Gang Xu 0001, Ran Ling, Renshu Gu
Comput. Aided Geom. Des.2
2024 LSOIT: Lexicon and Syntax Enhanced Opinion Induction Tree for Aspect-based Sentiment Analysis
Haiyan Wu, Di Zhou 0009, Chaoqun Sun, Zhiqiang Zhang 0010, Yong Ding 0003
Expert Syst. Appl.1
2024 Towards Human-Compatible Autonomous Car: A Study of Non-Verbal Turing Test in Automated Driving With Affective Transition Modelling
abstract
Autonomous cars are indispensable when humans go further down the hands-free route. Although existing literature highlights that the acceptance of the autonomous car will increase if it drives in a human-like manner, sparse research offers the naturalistic experience from a passenger's seat perspective to examine the humanness of current autonomous cars. The present study tested whether the AI driver could create a human-like ride experience for passengers based on 69 participants' feedback in a real-road scenario. We designed a ride experience-based version of the non-verbal Turing test for automated driving. Participants rode in autonomous cars (driven by either human or AI drivers) as a passenger and judged whether the driver was human or AI. The AI driver failed to pass our test because passengers detected the AI driver above chance. In contrast, when the human driver drove the car, the passengers' judgement was around chance. We further investigated how human passengers ascribe humanness in our test. Based on Lewin's field theory, we advanced a computational model combining signal detection theory with pre-trained language models to predict passengers' humanness rating behaviour. We employed affective transition between pre-study baseline emotions and corresponding post-stage emotions as the signal strength of our model. Results showed that the passengers' ascription of humanness would increase with the greater affective transition. Our study suggested an important role of affective transition in passengers' ascription of humanness, which might become a future direction for autonomous driving.
Zhaoning Li, Qiaoli Jiang, Zhengming Wu, Anqi Liu 0001, Haiyan Wu, Miner Huang, Kai Huang 0001, Yixuan Ku
IEEE Trans. Affect. Comput.5
2023 Towards human-compatible autonomous car: A study of non-verbal Turing test in automated driving with affective transition modelling
Zhaoning Li, Qiaoli Jiang, Zhengming Wu, Haiyan Wu, Miner Huang, Yixuan Ku
CogSci5
2023 Design Cloud-Edge Collaborated Batch Control Systems Based on Automatic Mapping IEC 61499 and ISA-88
abstract
Manufacturing is entering a new era with Industrial Internet and edge computing. The Industrial Internet cloud platform provides massive computing power and storage spaces for field devices. With more powerful chips available, field devices are also capable of handling multiple complex computational tasks simultaneously. How to collaborate resources from both cloud platforms and edge devices become an important topic for manufacturers. In this paper, a cloud-edge collaborated batch control system is proposed based on the IEC 61499 standard and the ISA-88 standard. The ISA-88 models are implemented as an independent IEC 61499 resource to support low-code development for batch control systems. Also, cloud resources are introduced in the IEC 61499 deployment to enable cloud-edge collaboration. Finally, the design process is verified with a liquid food processing line.
Jinbo Zhu, Weimin Lyu, Wenbin Dai, Haiyan Wu
IECON5
2023 DiagVol: Multi-block Bézier Volume Modeling from Prescribed Diagonal Surface Pairs
Qinghua Hu, Gang Xu 0001, Haiyan Wu, Yufei Pang
Comput. Aided Des.5
2023 Textile image recoloring by polarization observation
Haipeng Luan, Masahiro Toyoura, Renshu Gu, Takamasa Terada, Haiyan Wu, Takuya Funatomi, Gang Xu 0001
Vis. Comput.5
2022 Self-supervised Models are Good Teaching Assistants for Vision Transformers
abstract
Transformers have shown remarkable progress on computer vision tasks in the past year. Compared to their CNN counterparts, transformers usually need the help of distillation to achieve comparable results on middle or small sized datasets. Meanwhile, recent researches discover that when transformers are trained with supervised and self-supervised manner respectively, the captured patterns are quite different both qualitatively and quantitatively. These findings motivate us to introduce an self-supervised teaching assistant (SSTA) besides the commonly used supervised teacher to improve the performance of transformers. Specifically, we propose a head-level knowledge distillation method that selects the most important head of the supervised teacher and self-supervised teaching assistant, and let the student mimic the attention distribution of these two heads, so as to make the student focus on the relationship between tokens deemed by the teacher and the teacher assistant. Extensive experiments verify the effectiveness of SSTA and demonstrate that the proposed SSTA is a good compensation to the supervised teacher. Meanwhile, some analytical experiments towards multiple perspectives (e.g. prediction, shape bias, robustness, and transferability to downstream tasks) with supervised teachers, self-supervised teaching assistants and students are inductive and may inspire future researches.
Haiyan Wu, Yinqi Zhang, Shaohui Lin, Yuan Xie 0006, Xing Sun 0001, Ke Li 0015
ICML1
2022 Ergodic theorems for capacity preserving Z+d-actions
Haiyan Wu
Int. J. Approx. Reason.1
2022 Kuramoto Model-Based Analysis Reveals Oxytocin Effects on Brain Network Dynamics
abstract
The oxytocin effects on large-scale brain networks such as Default Mode Network (DMN) and Frontoparietal Network (FPN) have been largely studied using fMRI data. However, these studies are mainly based on the statistical correlation or Bayesian causality inference, lacking interpretability at the physical and neuroscience level. Here, we propose a physics-based framework of the Kuramoto model to investigate oxytocin effects on the phase dynamic neural coupling in DMN and FPN. Testing on fMRI data of 59 participants administrated with either oxytocin or placebo, we demonstrate that oxytocin changes the topology of brain communities in DMN and FPN, leading to higher synchronization in the FPN and lower synchronization in the DMN, as well as a higher variance of the coupling strength within the DMN and more flexible coupling patterns at group level. These results together indicate that oxytocin may increase the ability to overcome the corresponding internal oscillation dispersion and support the flexibility in neural synchrony in various social contexts, providing new evidence for explaining the oxytocin modulated social behaviors. Our proposed Kuramoto model-based framework can be a potential tool in network neuroscience and offers physical and neural insights into phase dynamics of the brain.
Shuhan Zheng, Zhichao Liang, Youzhi Qu, Qingyuan Wu, Haiyan Wu, Quanying Liu
Int. J. Neural Syst.5
2022 Phrase dependency relational graph attention network for Aspect-based Sentiment Analysis
Haiyan Wu, Zhiqiang Zhang 0010, Shaoyun Shi, Haiyu Song 0001
Knowl. Based Syst.1
2022 Symbolic sequence representation with Markovian state optimization
Lifei Chen, Haiyan Wu, Wenxuan Kang, Shengrui Wang
Pattern Recognit.2
2021 Contrastive Learning for Compact Single Image Dehazing
abstract
Single image dehazing is a challenging ill-posed problem due to the severe information degeneration. However, existing deep learning based dehazing methods only adopt clear images as positive samples to guide the training of dehazing network while negative information is unexploited. Moreover, most of them focus on strengthening the dehazing network with an increase of depth and width, leading to a significant requirement of computation and memory. In this paper, we propose a novel contrastive regularization (CR) built upon contrastive learning to exploit both the information of hazy images and clear images as negative and positive samples, respectively. CR ensures that the restored image is pulled to closer to the clear image and pushed to far away from the hazy image in the representation space.Furthermore, considering trade-off between performance and memory storage, we develop a compact dehazing network based on autoencoder-like (AE) framework. It involves an adaptive mixup operation and a dynamic feature enhancement module, which can benefit from preserving information flow adaptively and expanding the receptive field to improve the network’s transformation capability, respectively. We term our dehazing network with autoencoder and contrastive regularization as AECR-Net. The extensive experiments on synthetic and real-world datasets demonstrate that our AECR-Net surpass the state-of-the-art approaches. The code is released in https://github.com/GlassyWu/AECR-Net.
Haiyan Wu, Yanyun Qu, Shaohui Lin, Ruizhi Qiao, Zhizhong Zhang 0001, Yuan Xie 0006, Lizhuang Ma
CVPR1
2021 A Novel Convolutional Neural Network Model to Remove Muscle Artifacts from EEG
abstract
The recorded electroencephalography (EEG) signals are usually contaminated by many artifacts. In recent years, deep learning models have been used for denoising of electroencephalography (EEG) data and provided comparable performance with that of traditional techniques. However, the performance of the existing networks in electromyograph (EMG) artifact removal was limited and suffered from the over-fitting problem. Here we introduce a novel convolutional neural network (CNN) with gradually ascending feature dimensions and downsampling in time series for removing muscle artifacts in EEG data. Compared with other types of convolutional networks, this model largely eliminates the over-fitting and significantly outperforms four benchmark networks in EEGdenoiseNet. Our study suggested that the deep network architecture might help avoid overfitting and better remove EMG artifacts in EEG.
Chen Wei 0006, Mingqi Zhao, Quanying Liu, Haiyan Wu
ICASSP5
2021 Towards Compact Single Image Super-Resolution via Contrastive Self-distillation
abstract
Convolutional neural networks (CNNs) are highly successful for super-resolution (SR) but often require sophisticated architectures with heavy memory cost and computational overhead significantly restricts their practical deployments on resource-limited devices. In this paper, we proposed a novel contrastive self-distillation (CSD) framework to simultaneously compress and accelerate various off-the-shelf SR models. In particular, a channel-splitting super-resolution network can first be constructed from a target teacher network as a compact student network. Then, we propose a novel contrastive loss to improve the quality of SR images and PSNR/SSIM via explicit knowledge transfer. Extensive experiments demonstrate that the proposed CSD scheme effectively compresses and accelerates several standard SR models such as EDSR, RCAN and CARN. Code is available at https://github.com/Booooooooooo/CSD.
Yanbo Wang 0003, Shaohui Lin, Yanyun Qu, Haiyan Wu, Zhizhong Zhang 0001, Yuan Xie 0006, Angela Yao
IJCAI4
2021 Author Name Disambiguation Using Multiple Graph Attention Networks
abstract
The ambiguity of name entities is a common problem in information retrieval, which leads to the decline of retrieval quality. This makes name disambiguation particularly important. In academic field, the rapidly increasing large-scale of publications has imposed more challenges to the name disambiguation problem. Existing works mainly focus on leveraging content information to distinguish different name entities. In this paper, we consider jointly utilizing both content information and relational information to disambiguate the same name. Firstly, we construct a Heterogeneous Academic Network based on meta information of publications such as collaborators, institutions and venues. Then, we transform the network into separate homogeneous graphs. After that, we propose Graph Attention Networks to jointly learn content and relational information by optimizing an embedding vector. Finally, a clustering algorithm is presented to gather author names most likely representing the same person. The experiments show that our method is effective and outperforms the state-of-the-art methods in both precision and recall metrics.
Zhiqiang Zhang 0010, Chunqi Wu, Zhao Li 0007, Juanjuan Peng, Haiyan Wu, Haiyu Song 0001, Shengchun Deng
IJCNN5
2021 Research on the Difference Between the Internet of Things and the Traditional Internet Under Artificial Intelligence
abstract
The Internet of Things is the most heated topic in the current electronic information technology industry. It is generally believed that it will become a new driving force for the development of human industry, and its comprehensive application will promote the human society to step forward to the “smart earth”. In this paper, based on artificial intelligence and Internet of Things technology, experimental design of fuzzy control and PID control algorithm anti-interference performance test simulation box. Experimental data show that due to the Internet of Things control system in the process of application, the controlled object in different degrees of non-linear, large lag, parameter time-varying and model uncertainty characteristics. It can be seen from the principle and simulation results that fuzzy control does not require the precise model of the controlled object and has strong adaptability. The experimental results show that if the pulses with a height of 0.1, a period of 10s, a width of 2S and a delay of 4s are set, interference is added to the output of a sampling time controller. Among them, T1=2, T2=1, =0.5, then the initial PID Kp0=7.6, Ki0=1.12, Kd0=4.88 in the fuzzy PID control system. The Internet of Things control system server has the advantages of comprehensive functions, good expansibility, easy operation and maintenance, etc., and the adoption of automated testing technology improves the efficiency and standardization of the server design process.
Yongjun Qi, Haiyan Wu
IWCMC2
2021 Application of Computer Software Processing Technology in Performance Information Management System
abstract
In the urgent need of large-scale research, the use of high-performance computer processing technology for performance information management has become a popular trend. This paper aims to evaluate the current situation of performance management and put forward further research direction. The expected improvements in performance, accountability, transparency, service quality and cost-effectiveness of performance information management have not yet been achieved. There are three problems in performance management: technical problems, institutional problems and participation problems. In order to build a comprehensive performance management system, external imposed restructuring and restructuring limit the successful implementation of performance management. This paper introduces the architecture and prototype implementation of a web service performance management system based on cluster. The system supports multiple types of web service traffic, dynamically allocates server resources, and maximizes the expected value of a given cluster utility function in the case of load fluctuations. The cluster utility is a performance function that provides each class with different services. The results show that the weight of performance appraisal index is the highest, accounting for 31%.
Haiyan Wu, Yongjun Qi
IWCMC1
2020 Modularized Syntactic Neural Networks for Sentence Classification
abstract
This paper focuses on tree-based modeling for the sentence classification task.In existing works, aggregating on a syntax tree usually considers local information of sub-trees.In contrast, in addition to the local information, our proposed Modularized Syntactic Neural Network (MSNN) utilizes the syntax category labels and takes advantage of the global context while modeling sub-trees.In MSNN, each node of a syntax tree is modeled by a label-related syntax module.Each syntax module aggregates the outputs of lower-level modules, and finally, the root module provides the sentence representation.We design a tree-parallel mini-batch strategy for efficient training and predicting.Experimental results on four benchmark datasets show that our MSNN significantly outperforms previous state-of-the-art tree-based methods on the sentence classification task.
Haiyan Wu, Shaoyun Shi
EMNLP (1)1
2020 Discriminative Clip Mining for Video Anomaly Detection
abstract
Real-world anomalous events are complicated, diverse, and rarely occurred. The main challenge to anomaly detection is to learn normal and anomalous patterns accurately. In this work, we propose discriminative clip mining for anomaly detection and classification: firstly, by introducing clip-level class activation mapping, an efficient Discriminative Anomalous Clip Miner (DACM) is developed to mine discriminative anomalous clips from a large number of normal ones; secondly, with the mined discriminative clips, an attentive ranking loss is designed to increase the anomalous instances hit rate of the traditional Multiple Instance Learning (MIL) model. Furthermore, by integrating the DACM and attentive MIL, one novel anomaly detection framework is proposed to learn more contrastive anomalous and normal patterns, and thus higher recognition performance can be achieved. Our experimental results on the widely-used UCF-Crime dataset show that, as compared to the state-of-the-art approaches, the proposed method achieves competitive performance both in anomaly detection and anomalous activity classification.
Wu Luo, Haiyan Wu
ICIP4
2020 CCAE: Cross-field categorical attributes embedding for cancer clinical endpoint prediction
Youru Li, Zhenfeng Zhu, Haiyan Wu, Silu Ding, Yao Zhao 0001
Artif. Intell. Medicine3
2019 Improving Action Recognition with the Graph-Neural-Network-based Interaction Reasoning
abstract
Recent human action recognition methods mainly model a two-stream or 3D convolution deep learning network, with which humans spatial-temporal features can be exploited and utilized effectively. However, due to the ignoring of interaction exploiting, most of these methods cannot get good enough performance. In this paper, we propose a novel action recognition framework with Graph Convolutional Network (GCN) based Interaction Reasoning: Objects and discriminative scene patches are detected using an object detector and class active mapping (CAM), respectively; and then a GCN is introduced to model the interaction among the detected objects and scene patches. Evaluation of two widely used video action benchmarks shows that the proposed work can achieve comparable performance: the accuracy up to 43.6% at EPIC Kitchen, and 47.0% at VLOG benchmark without using optical flow, respectively.
Wu Luo, Haiyan Wu
VCIP4
2018 Agglomeration Detection in Gas-Phase Ethylene Polymerization Based on Multi-scale Convolutional Neural Network
Jing Wang 0016, Haiyan Wu
ICONIP (4)3
2018 A Neurorobotic Experiment for Crossmodal Conflict Resolution in Complex Environments
abstract
Crossmodal conflict resolution is crucial for robot sensorimotor coupling through the interaction with the environment, yielding swift and robust behaviour also in noisy conditions. In this paper, we propose a neurorobotic experiment in which an iCub robot exhibits human-like responses in a complex crossmodal environment. To better understand how humans deal with multisensory conflicts, we conducted a behavioural study exposing 33 subjects to congruent and incongruent dynamic audio-visual cues. In contrast to previous studies using simplified stimuli, we designed a scenario with four animated avatars and observed that the magnitude and extension of the visual bias are related to the semantics embedded in the scene, i.e., visual cues that are congruent with environmental statistics (moving lips and vocalization) induce the strongest bias. We implement a deep learning model that processes stereophonic sound, facial features, and body motion to trigger a discrete behavioural response. After training the model, we exposed the iCub to the same experimental conditions as the human subjects, showing that the robot can replicate similar responses in real time. Our interdisciplinary work provides important insights into how crossmodal conflict resolution can be modelled in robots and introduces future research directions for the efficient combination of sensory observations with internally generated knowledge and expectations.
German Ignacio Parisi, Pablo V. A. Barros, Di Fu, Sven Magg, Haiyan Wu, Xun Liu 0001, Stefan Wermter
IROS5
2018 Fuzzy clustering based pseudo-swept volume decomposition for hexahedral meshing
Haiyan Wu, Shuming Gao, Rui Wang 0004
Comput. Aided Des.1
2017 Sheet operation based block decomposition of solid models for hex meshing
Rui Wang 0004, Haiyan Wu, Shuming Gao
Comput. Aided Des.4
2017 Unified Architecture of Active Fault Detection and Partial Active Fault-Tolerant Control for Incipient Faults
abstract
Incipient faults are difficult to be detected due to the intrinsic fault tolerance of traditional controller, but it should be eliminated as soon as possible before it deteriorates with time into something more serious. As a consequence of an intrinsic inability to assess whether a fault occurs based on output residual, the existing detection methods are failure for incipient fault. So the aim of active fault detection (AFD) is to make the system be unstable when incipient fault has occurred, which drives rapid fault detection. The fault-tolerant control (FTC) is designed to maintain the system stable and eliminate the fault impact without shutting the process down even if faults occur. In this paper, the Youla-Jabr-Bongiorno-Kucera (YJBK) parameter is employed to build the AFD and the FTC based on the relationship analysis between the fault and the dual YJBK parameter. A new structure of the tolerant controller parameter for FTC is designed, named as partial active FTC (PAFTC). PAFTC is dependent upon the fault detection information but not the fault size considering the parameter fault with unknown size and known form. A unified operation architecture for AFD and PAFTC with different YJBK parameters for incipient faults is proposed. Some illustrative examples are given to indicate the effectiveness of the proposed unified operation architecture.
Jing Wang 0016, Bo Qu, Haiyan Wu
IEEE Trans. Syst. Man Cybern. Syst.4
2016 Visual servoing for object manipulation: A case study in slaughterhouse
abstract
Automation for slaughterhouse challenges the design of the control system due to the variety of the objects. Realtime sensing provides instantaneous information about each piece of work and thus, is useful for robotic system developed for slaughterhouse. In this work, a pick and place task which is a common task among tasks in slaughterhouse is selected as the scenario for the system demonstration. A vision system is utilized to grab the current information of the object, including position and orientation. The information about the object is then transferred to the robot side for path planning. An online and offline combined path planning algorithm is proposed to generate the desired path for the robot control. An industrial robot arm is applied to execute the path. The system is implemented for a lab-scale experiment, and the results show a high success rate of object manipulation in the pick and place task. The approach is implemented in ROS which allows utilization of the developed algorithm on different platforms with little extra effort.
Haiyan Wu, Thomas Timm Andersen, Nils A. Andersen, Ole Ravn
ICARCV1
2016 An approach to achieving optimized complex sheet inflation under constraints
Shuming Gao, Rui Wang 0004, Haiyan Wu
Comput. Graph.4
2012 Ping-pong robotics with high-speed vision system
abstract
The performance of vision-based control is usually limited by the low sampling rate of the visual feedback. We address Ping-Pong robotics as a widely studied example which requires high-speed vision for highly dynamic motion control. In order to detect a flying ball accurately and robustly, a multi-threshold segmentation algorithm is applied in a stereo-vision running at 150Hz. Based on the estimated 3D ball positions, a novel two-phase trajectory prediction is exploited to determine the hitting position. Benefiting from the high-speed visual feedback, the hitting position and thus the motion planning of the manipulator are updated iteratively with decreasing error. Experiments are conducted on a 7 degrees of freedom humanoid robot arm. A successful Ping-Pong playing between the robot arm and human is achieved with a high successful rate of 88%.
Hailing Li, Haiyan Wu, Lei Lou, Kolja Kühnlenz, Ole Ravn
ICARCV2
2011 Performance-oriented networked visual servo control with sending rate scheduling
abstract
In order to speed up image processing in visual servoing, the distributed computational power across networks and appropriate data transmission mechanisms are of particular interest. In this paper, a high sampling rate of visual feedback is achieved by distributed computation on a cloud image processing platform. For target tracking with a networked visual servo control system, a switching control law considering the varying feedback delay caused by image processing and data transmission is applied to improve the control performance. A sending rate scheduling strategy aiming at saving the network load is proposed based on the tracking error. Experiments on a 7 degree-of-freedom (DoF) manipulator are carried out to validate the proposed approach. The proposed approach shows a similar control performance as a system without sending rate scheduling, however, beneficially with largely reduced network load.
Haiyan Wu, Lei Lou, Chih-Chung Chen, Sandra Hirche, Kolja Kühnlenz
ICRA1
2010 A framework of networked visual servo control system with distributed computation
abstract
In this paper, a networked visual servo control system with distributed computation is proposed to overcome the low sampling rate problem in vision-based control systems. A real-time image data transmission protocol based on Realtime Transport Protocol (RTP) is developed. The captured images are sent to different processing nodes connected over a communication network and processed in parallel. Thus, a high sampling rate of the visual feedback is achieved under a cloud image processing architecture. The varying image processing delay caused by the varying number of extracted features and the random transmission delay are modeled as a random process with Bernoulli distribution. By using the input-delay approach, the resulted networked visual servo control system is reformulated into a stochastic continuous-time system with time-varying delay. Experiments on two 1-DoF linear motor modules are carried out to validate the proposed approach. A visual servo control system without parallel distributed computation is implemented for comparison. The experimental results demonstrate significant performance improvement by the proposed approach.
Haiyan Wu, Lei Lou, Chih-Chung Chen, Sandra Hirche, Kolja Kühnlenz
ICARCV1
2010 A switching control law for a networked visual servo control system
abstract
In this paper, a novel switching controller is proposed for a networked visual servo control system with varying feedback delay due to image processing and data transmission. The varying image processing delay caused by the varying number of extracted features for pose estimation due to different view angles, illumination conditions and noise, is modeled by its occurrence probability. The time delay due to transmission over the communication network is also modeled as random process. By using a sampled-data system approach and an input-delay approach, the linearized visual servo control system is reformulated into a stochastic continuous-time system with time-varying delay. A novel stability condition and associated switching controller are derived based on the occurrence probabilities of delays. Experiments on a 1-DoF linear module equipped with a camera are conducted to validate the proposed approach. A non-switching controller approach is implemented for comparison. The experimental results demonstrate significant performance improvement of the proposed control approach.
Haiyan Wu, Chih-Chung Chen, Jiayun Feng, Kolja Kühnlenz, Sandra Hirche
ICRA1
2010 Distributed computation and data scheduling for networked visual servo control systems
abstract
The stability and performance of visual servo control systems strongly depend on the delays caused by image processing. In order to accelerate the visual feedback, the distributed computational power across networks and appropriate data transmission mechanism are of particular interest. In this paper, a novel distributed computation with data scheduling is proposed for networked visual servo control systems (NVSCSs) aiming at improving the control performance. A realtime transport protocol is developed for image data transmission. For a NVSCS which is modeled as a continuous-time system with computation, transmission and holding delays, a switching control law is applied. A probabilistic sampling scheduler is derived such that the control performance and the network load caused by image data transmission are balanced. Experiments on two 1-DoF linear modules equipped with a camera are conducted to validate the proposed approach. A visual servo system without data scheduling is implemented for comparison. The experimental results demonstrate a comparable control performance of the proposed approach with an advantage of reduced network load.
Haiyan Wu, Lei Lou, Chih-Chung Chen, Kolja Kühnlenz, Sandra Hirche
IROS1
2009 An explorative study of visual servo control with insect-inspired Reichardt-model
abstract
In this paper, an insect-inspired motion detector (Reichardt-model) is applied to visual servo control to ensure the stability of the system with high gain and time delay in its feedback. A Reichardt-based control scheme is compared with a conventional visual servoing approach. As a consequence of the specific velocity dependence of the Reichardt-model, the stability margin of the visual servo control is increased and high overall gains, thus, better performance are achievable. The response of the Reichardt-model in the experiment and the control performance of velocity control approach with the Reichardt-model in the closed loop are investigated. The velocity control model is tested on a 1-DOF linear motor module with different feedback gain and different time delay in the loop. The results of simulation and realtime experiments demonstrate the stabilizing character of the Reichardt-based approach.
Haiyan Wu, Tianguang Zhang, Alexander Borst, Kolja Kühnlenz, Martin Buss
ICRA1
2008 An Automatic Scheme to Categorize User Sessions in Modern HTTP Traffic
abstract
The characterization of HTTP traffic is crucial for performance evaluation and server design. In this paper, we analyze massive Web traces generated by various busy servers in recent years, trying to find the new features of modern HTTP traffic and user behaviors. Comparing the conclusions of earlier studies with our results, we have spotted considerable unconventional ingredients in modern HTTP traffic that could hardly be described by previous models. We also propose an innovative scheme to automatically categorize these various ingredients in modern traffic. The novel aspects of our work are: (1)It reveals the sophisticated composition of modern HTTP traffic with solid evidence, (2)It provides an automatic method to analyze the composition of modern HTTP traffic and (3)It promises a powerful manner to evaluate the possible performance implication of modern HTTP traffic on existing Web servers. We hope this work would help researchers and designers to better understand new features of HTTP workloads and therefore make corresponding adaptations in design practice.
Xiaozhu Lin, Lin Quan, Haiyan Wu
GLOBECOM3
2008 An FPGA implementation of insect-inspired motion detector for high-speed vision systems
abstract
In this paper, an array of biologically inspired elementary motion detectors (EMDs) is implemented on an FPGA (Field Programmable Gate Array) platform. The well-known Reichardt-type EMD, modeling the insect’s visual signal processing system, is very sensitive to motion direction and has low computational cost. A modified structure of EMD is used to detect local optical flow. Six templates of receptive fields, according to the fly’s vision system, are designed for simple ego-motion estimation. The results of several typical experiments demonstrate local detection of optical flow and simple motion estimation under specific backgrounds. The performance of the real-time implementation is sufficient to deal with a video frame rate of 350 fps at 256 x 256 pixels resolution. The execution of the motion detection algorithm and the resulting time delay is only 0.25 μs. This hardware is suited for obstacle detection, motion estimation and UAV/MAV attitude control.
Tianguang Zhang, Haiyan Wu, Alexander Borst, Kolja Kühnlenz, Martin Buss
ICRA2
2006 Double Inverted Pendulum Control Based on Support Vector Machines and Fuzzy Inference
Han Liu 0007, Haiyan Wu, Fucai Qian
ISNN (2)2