EDBT 2026 Demo / reviewers in the wild / expert
Hongyu Yang 0002
dblp:57/5473-2
· DBLP profile ↗
62ranked-venue papers
0as first author
50since 2021 · last 2026
0000-0002-2322-0538ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 18 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Joint scheduling of runway and taxiway considering uncertain taxiing time based on an improved ACO algorithm
Lirong Zhang, Yi Lin 0006, Suwan Yin, Wu Deng 0001, Hongyu Yang 0002, Jianwei Zhang 0013 |
Inf. Sci. | 5 |
| 2026 | DSTFGCN: A dynamic spatial-temporal fusion graph convolution network for traffic flow forecastingabstractTraffic flow prediction is one of the core technology of Intelligent Transportation System. Its fundamental challenge is to effectively model the complex spatial-temporal dependencies. Although extensive research has been conducted in this field, the limitations of current methods restrict their effectiveness in accurate predictions. For temporal dependence, existing methods based on recurrent neural networks only focus on local dependencies and ignore global dependencies. For spatial dependencies, existing methods use predefined or adaptive adjacency matrices that cannot accurately reflect the relationships between real traffic flow. To overcome these limitations, we propose a dynamic spatial-temporal fusion graph convolution network (DSTFGCN). In the temporal aspect, we introduce gated dilated causal convolution to capture the local dependencies and node-independent temporal graph convolution to capture the global dependencies specific to each node. In the spatial aspect, we propose a dynamic graph convolution block. It can construct dynamic graphs based on the characteristics of the input data and aggregate both local and global spatial dependencies. Experiments on six real-world datasets have shown that DSTFGCN outperforms current mainstream methods. The codes are available at https://github.com/SYLan2019/DSTFGCN. Tianyi Pan, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Hongyu Yang 0002, Zhiang Hou, Yao Ren |
Neural Networks | 5 |
| 2026 | PaTH: Patch-wise temporal hierarchical modeling for high-resolution flight trajectory prediction
Guoxin Huang, Fengshuo Ye, Yi Lin 0006, Hongyu Yang 0002, Yunxiang Han, Dongyue Guo |
Pattern Recognit. | 4 |
| 2026 | Physics-Guided Neural Radiance Fields for Forward-Looking Sonar ImagingabstractWhile neural radiance fields (NeRFs) have achieved remarkable success in optical imaging, their extension to forwardlooking sonar (FLS) remains underexplored due to the fundamental mechanisms of different acoustic propagation physics. In this work, we present Sonar-NeRF, a physics-guided neural rendering framework tailored for high-fidelity FLS novel view synthesis. This approach replaces volume rendering with an explicit forward rendering model, derived from the active sonar equation to directly predict sonar echo intensities. A differentiable acoustic reflection model is incorporated to effectively capture specular reflections on metallic surfaces. In addition, heteroscedastic uncertainty learning based on an additive-multiplicative noise model is introduced to enable adaptive noise modelling. Experiments on synthetic and real FLS data show that our method is quantitatively and qualitatively superior to the conventional sonar simulators and existing NeRF-based methods. Cao Huang, Jinchang Ren, Hongyu Yang 0002, Yulong Ji |
IEEE Signal Process. Lett. | 3 |
| 2026 | Integrated Dynamic Routing and Off-Block Optimization Based on Taxiing Route Impedance for Airport Surface OperationsabstractThis study develops an integrated dynamic routing and off-block (IDRO) optimization framework for real-time airport surface operations. Instead of relying on complex taxi time prediction models, the proposed approach employs adynamic route impedanceformulation as a practical heuristic to capture real-time congestion conditions on the airport surface. The impedance integrates domain-specific operational factors, including free flow taxi time, turning penalties, and potential aircraft conflicts, enabling real-time route evaluation and decision-making without extensive simulation. Building upon this impedance representation, the IDRO framework coordinates both taxi routing and off-block release timing under a real-time First-Come-First-Served (FCFS) principle. The approach was implemented and tested using a high-fidelity cellular automata simulator calibrated with operational data from Beijing Capital International Airport. Compared with real-world baseline operations, the proposed framework reduced average taxi time, flight delay, and taxi conflicts by 10.43%, 24.27%, and 31.82%, respectively. A comparative study with another similar approach in the literature also reveals the advantage of the proposed strategy. These results demonstrate the framework’s computational efficiency, consistent performance across scenarios, and strong potential for real-world deployment in large-scale airport environments. Suwan Yin, Washington Yotto Ochieng, Hongyu Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2026 | AutoFPDesigner: Automated Flight Procedure Design Based on Multi-Agent Large Language ModelabstractFlight procedures are essential to the safety and efficiency of air traffic management. However, due to the highly specialized nature of the flight procedure design process, existing methods rely heavily on manual operations and adjustments with limited automation, resulting in inefficiencies and potential safety risks. This study introduces AutoFPDesigner, a new agent-driven approach to flight procedure design, leveraging large language models (LLMs). By utilizing multi-agent collaboration, AutoFPDesigner automates Performance-Based Navigation (PBN) procedures, enabling end-to-end automation. In this framework, the designer’s role shifts from an executor to a supervisor, issuing tasks through natural language, while the system integrates specialized knowledge and uses a toolset to complete the design. Experimental results show that procedures designed with this approach meet safety requirements nearly 100%, with 75% of tasks completed in a limited number of steps. Moreover, AutoFPDesigner performs effectively across various design tasks, outperforming existing methods. Additionally, this study conducted human interaction experiments and introduced an “instruction-based” feedback method to address agent misinterpretation of human feedback. Experimental results demonstrate that the system bridges the skill gap between experts and beginners, and that the “instruction-based” feedback method enhances the accuracy of agent feedback interpretation. Code and data are available onhttps://github.com/Zhulongtao6/AutoFPDesigner-LLM Longtao Zhu, Hongyu Yang 0002, Yulong Ji, Jinchang Ren |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | RAV: Retrieval-Augmented Voting for Tactile Descriptions Without TrainingabstractTactile perception is essential for humanenvironment interaction, and deriving tactile descriptions from multimodal data is a key challenge for embodied intelligence to understand human perception.Conventional approaches relying on extensive parameter learning for multimodal perception are rigid and computationally inefficient.To address this, we introduce Retrieval-Augmented Voting (RAV), a parameter-free method that constructs visualtactile cross-modal knowledge directly.RAV retrieves similar visual-tactile data for given visual and tactile inputs and generates tactile descriptions through a voting mechanism.In experiments, we applied three voting strategies, SyncVote, DualVote and WeightVote, achieving performance comparable to large-scale crossmodal models without training.Comparative experiments across datasets of varying quality-defined by annotation accuracy and data diversity-demonstrate that RAV's performance improves with higher-quality data at no additional computational cost.Code, and model checkpoints are opensourced at https: //github.com/PluteW/RAV. Yulong Ji, Hongyu Yang 0002 |
EMNLP | 3 |
| 2025 | Learning Geometry-Aware Representation for Gaze EstimationabstractAppearance-based gaze estimation has achieved remarkable progress in recent years. However, the inherent geometry characteristics of eye and facial areas are not fully explored in existing methods, which limits the generalization and robustness of the model. In this paper, we propose a novel end-to-end framework for cross-domain gaze estimation by integrating latent geometric representation into appearance-based gaze framework. More specifically, we first exploit the 3DMM method to fit unconstrained faces and eyes in the wild, which would generate adaptive normal information with explicit 3D geometry prior. Then we joint the normal map and the corresponding RGB appearance information to infer the 3D gaze direction with carefully-designed spatial-frequency attention and local-global feature interaction modules. The key to our method is to integrate explicit 3D geometry representation into a 2D learning architecture, which leads to a better trade-off between performance and efficiency. Experiments on both MPIIGaze and EyeDiap datasets demonstrate that the proposed method achieves the state-of-the-art accuracy of 3.56° and 5.10° separately, and also presents superior generalization ability on cross-domain dataset evaluations. Qida Tan, Wenchao Du, Hu Chen 0002, Hongyu Yang 0002 |
ICIP | 5 |
| 2025 | EMGPose: An Efficient Multi-Granularity Representation for Human Pose EstimationabstractCurrent Transformer-based methods typically unwisely represent the entire image at a single granularity. A high-resolution representation of the regions of interest can significantly improve the accuracy of human pose estimation while causing unnecessary computational costs for other regions. To overcome this limitation, we propose an efficient two-stage framework using adaptive Multi-Granularity representation for different important image regions for Human pose estimation (EMGPose). In the first stage, the image is split into coarse-grained patches for simple inference. If without sufficient accuracy, important patches will be resplit into multiple finer-grained patches for the second stage of inference. Furthermore, we propose a new token-merge strategy based on token importance and similarity in Transformer, effectively reducing the computational load from low-information background patches. Extensive experiments demonstrate the excellent performance of the proposed method. Specifically, our model EMGPose-Base achieves 76.3 AP (+0.5 AP) and 62.2 AP (+2.6 AP) and higher efficiency than baseline ViTPose-Base on the COCO validation set and OCHuman test set, respectively. Guonan Deng, Shiyong Lan, Wenwu Wang 0001, Yixin Qiao, Yao Li 0019, Haohan Chen, Hongyu Yang 0002 |
ICME | 7 |
| 2025 | FE-HGAT: Frequency-Enhanced Hybrid Graph Attention Network For Traffic PredictionabstractTraffic flow prediction is crucial for urban traffic management and planning. However, although existing research has achieved promising results, most methods primarily focus on time-domain processing, with insufficient exploration of frequency-domain signal characteristics. Moreover, existing approaches often fail to effectively distinguish and simultaneously model the spatial dependencies between nearby and distant nodes. To address these issues, this paper proposes a Frequency-Enhanced Hybrid Graph Attention Network (FE-HGAT) for traffic flow prediction. Our approach employs a dynamic filter in the temporal dimension, which utilizes Fast Fourier Transform (FFT) to enhance key frequency-domain features, thereby better characterizing the temporal dependencies in traffic data. Besides, to capture spatial dependencies between nodes at varying distances, we design a dynamic threshold module to distinguish between nearby and distant nodes, employing external attention (EA) and a mixture-of-experts-enhanced graph attention network (MOE-GAT) to model local dependencies and long-distance semantic similarities, respectively. Experiments demonstrate that FE-HGAT outperforms existing baseline models on several public transportation datasets, validating its effectiveness in traffic forecasting. The code is available at https://github.com/ry123scuer/FE-HGAT. Yao Ren, Wujiang Zhu, Shiyong Lan, Xinyuan Zhou, Hongyu Yang 0002, Zhiang Hou |
SMC | 5 |
| 2025 | Taming deep reinforcement learning-based conflict resolution in air traffic control using geometric technique
Hongyu Yang 0002, Yunxiang Han, Suwan Yin |
Expert Syst. Appl. | 2 |
| 2025 | Enhancing air traffic control: A transparent deep reinforcement learning framework for autonomous conflict resolution
Hongyu Yang 0002, Yi Lin 0006, Suwan Yin |
Expert Syst. Appl. | 2 |
| 2025 | FAcupoint: The first dense facial acupoint localization dataset and baselines
Jizhe Zhou 0001, Hongyu Yang 0002, Yi Lin 0006 |
Expert Syst. Appl. | 4 |
| 2025 | Geometric self-supervision for monocular 3D animal pose estimation
Xiaowei Dai, Shuiwang Li, Qijun Zhao, Hongyu Yang 0002 |
Pattern Recognit. | 4 |
| 2025 | Geometry-Aware Appearance Learning for Generalized Gaze Estimation
Qida Tan, Wenchao Du, Hu Chen 0002, Hongyu Yang 0002 |
IEEE Signal Process. Lett. | 5 |
| 2025 | A Non-Autoregressive Multi-Horizon Flight Trajectory Prediction Framework With Gray Code RepresentationabstractFlight Trajectory Prediction (FTP) is an essential task in Air Traffic Control (ATC), which can assist air traffic controllers in managing airspace more safely and efficiently. Existing methods generally perform multi-horizon FTP tasks in an autoregressive manner, thereby suffering from error accumulation and low-efficiency problems. In this paper, a novel framework, called FlightBERT++, is proposed to i) forecast multi-horizon flight trajectories directly in a non-autoregressive way, and ii) improve the limitation of the binary encoding (BE) representation in the FlightBERT framework. Specifically, the proposed framework is implemented by a generalized encoder-decoder architecture, in which the encoder learns the temporal-spatial patterns from historical observations and the decoder predicts the flight status for the future horizons. Compared to conventional architecture, an innovative horizon-aware context generator is dedicatedly designed to consider the prior horizon information, which further enables non-autoregressive multi-horizon prediction. Additionally, the Gray code representation and the differential prediction paradigm are designed to cope with the high-bit misclassifications of the BE representation, which significantly reduces the outliers in the predictions. Moreover, a differential prompted decoder is proposed to enhance the capability of the differential predictions by leveraging the stationarity of the differential sequence. Extensive experiments are conducted to validate the proposed framework on a real-world flight trajectory dataset. The experimental results demonstrated that the proposed framework outperformed the competitive baselines in both FTP performance and computational efficiency. The code is publicly available at: https://github.com/gdy-scu/FlightBERT_PP_V2 Dongyue Guo, Fengshuo Ye, Jianwei Zhang 0013, Hongyu Yang 0002, Yi Lin 0006 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Solving Zero-Shot Sparse-View CT Reconstruction With Variational Score SolverabstractComputed tomography (CT) stands as a ubiquitous medical diagnostic tool. Nonetheless, the radiation-related concerns associated with CT scans have raised public apprehensions. Mitigating radiation dosage in CT imaging poses an inherent challenge as it inevitably compromises the fidelity of CT reconstructions, impacting diagnostic accuracy. While previous deep learning techniques have exhibited promise in enhancing CT reconstruction quality, they remain hindered by the reliance on paired data, which is arduous to procure. In this study, we present a novel approach named Variational Score Solver (VSS) for sparse-view reconstruction without paired data. Our approach entails the acquisition of a probability distribution from densely sampled CT reconstructions, employing a latent diffusion model. High-quality reconstruction outcomes are achieved through an iterative process, wherein the diffusion model serves as the prior term, subsequently integrated with the data consistency term. Notably, rather than directly employing the prior diffusion model, we distill prior knowledge by finding the fixed point of the diffusion model. This framework empowers us to exercise precise control over the process. Moreover, we depart from modeling the reconstruction outcomes as deterministic values, opting instead for a distribution-based approach. This enables us to achieve more accurate reconstructions utilizing a trainable model. Our approach introduces a fresh perspective to the realm of zero-shot CT reconstruction, circumventing the constraints of supervised learning. Extensive qualitative and quantitative experiments unequivocally demonstrate that VSS surpasses other contemporary unsupervised and achieves comparable results compared to the most advanced supervised methods in sparse-view reconstruction tasks. Codes are available in https://github.com/fpsandnoob/vss. Linchao He, Wenchao Du, Peixi Liao, Fenglei Fan, Hu Chen 0002, Hongyu Yang 0002, Yi Zhang 0018 |
IEEE Trans. Medical Imaging | 6 |
| 2024 | A Progressively Prompt-guided Model for Sparse-View CT ReconstructionabstractWhile sparse-view Computed Tomography (CT) has a remarkable impact on reducing ionizing radiation dose while accelerating data acquisition, the reconstructed images have been compromised by streak-like artifacts, affecting clinical diagnostics. By integrating powerful regularization with deep learning technologies into iterative reconstruction algorithms, the deep-unrolling-based methods have achieved promising results in terms of reconstruction quality and theoretical interpretability. However, leading methods always focus on learning powerful content priors with diverse technologies and ignoring the latent noise distribution prior in the image domain, thereby limiting the ability of structure-preserving and detail reconstructing of the model. To alleviate this problem, we propose a Progressively Prompt-guided Model (shorted by PPM) for sparse-view CT reconstruction. Specifically, we inject the idea of prompt learning into an iterative unrolled neural network, in which a learnable prompt module is inserted into each unrolled block to perceive image content and noise distribution in a self-adaptive manner, which leads to the more powerful priors to guide high-quality CT image reconstruction. Furthermore, we construct a progressively guiding strategy to facilitate high-quality prompt generation while speeding model convergence. Extensive experiments demonstrate that our PPM achieves state-of-the-art performance in artifact suppression, structure fidelity, and visual perception similarity. The code is available at https://github.com/Wenchao-Du/PPM/. Qiao Mu, Hu Chen 0002, Wenchao Du, Hongyu Yang 0002 |
BIBM | 6 |
| 2024 | Dtpose: Learning Disentangled Token Representation For Effective Human Pose EstimationabstractExploring rich visual clues and spatial geometric constraints to locate keypoints is essential for human pose estimation. Existing Transformer-based methods have presented unique advantages via token representation, where each keypoint is explicitly embedded as a token to learn visual appearance clues and geometric relationships simultaneously from images. However, it is difficult to learn powerful pose representation via self-attention mechanism due to latent interference, e.g., blurring and self-occlusion. To alleviate this challenge, we present a novel framework that Disentangles hybrid Token representation to explore more effective visual and keypoint information for Pose estimation (termed by DTPose). In detail, DTPose contains two key modules. First, the Disentangled Token Representation module is used to explore visual clues and geometry constraints sequentially, which alleviates the noise interference and enables the geometry and appearance clues to be exploited more sufficiently. Furthermore, the Hierarchical Spatial Decoding head is exploited to preserve the 2 D geometric structure information of keypoints as much as possible. Extensive experiments on COCO dataset demonstrate significant performance gains of our DTPose, which achieves 76.5 (+0.7) AP and 75.7 (+0.6) AP than the TokenPose-L on the COCO validation and test-dev sets separately. Shiyang Ye, Hu Chen 0002, Wenchao Du, Hongyu Yang 0002 |
ICIP | 6 |
| 2024 | Depth-Aware Dual-Stream Interactive Transformer Network for Facial Expression Recognition
Yiben Jiang, Xiao Yang 0029, Keren Fu, Hongyu Yang 0002 |
PRCV (11) | 4 |
| 2024 | Exploring Fast and Flexible Zero-Shot Low-Light Image/Video EnhancementabstractAbstract Low‐light image/video enhancement is a challenging task when images or video are captured under harsh lighting conditions. Existing methods mostly formulate this task as an image‐to‐image conversion task via supervised or unsupervised learning. However, such conversion methods require an extremely large amount of data for training, whether paired or unpaired. In addition, these methods are restricted to specific training data, making it difficult for the trained model to enhance other types of images or video. In this paper, we explore a novel, fast and flexible, zero‐shot, low‐light image or video enhancement framework. Without relying on prior training or relationships among neighboring frames, we are committed to estimating the illumination of the input image/frame by a well‐designed network. The proposed zero‐shot, low‐light image/video enhancement architecture includes illumination estimation and residual correction modules. The network architecture is very concise and does not require any paired or unpaired data during training, which allows low‐light enhancement to be performed with several simple iterations. Despite its simplicity, we show that the method is fast and generalizes well to diverse lighting conditions. Many experiments on various images and videos qualitatively and quantitatively demonstrate the advantages of our method over state‐of‐the‐art methods. Xianjun Han, Taoli Bao, Hongyu Yang 0002 |
Comput. Graph. Forum | 3 |
| 2024 | Hierarchical disentangled representation for image denoising and beyond
Wenchao Du, Hu Chen 0002, Yi Zhang 0018, Hongyu Yang 0002 |
Image Vis. Comput. | 4 |
| 2024 | Integrating prior knowledge into a bibranch pyramid network for medical image segmentation
Xianjun Han, Can Bai, Hongyu Yang 0002 |
Image Vis. Comput. | 4 |
| 2024 | LSTPNet: Long short-term perception network for dynamic facial expression recognition in the wild
Chengcheng Lu, Yiben Jiang, Keren Fu, Qijun Zhao, Hongyu Yang 0002 |
Image Vis. Comput. | 5 |
| 2024 | Long-Term Airport Network Performance Forecasting With Linear Diffusion Graph NetworksabstractPrecise forecasting of airport performances, such as landing rates and delays, is essential for the smooth operation of air traffic management systems and for improving the passenger experience. While current efforts predominantly address short-term predictions, the imperative for long-term forecasting is undeniable, particularly for strategic operational planning and resource management. Equally important is the explainability of these forecasts, which is critical for effective decision-making. To meet these needs, our study introduces an innovative approach to airport performance forecasting with the Linear-Diffusion Graph Network (LDGN), an explainable and probabilistic model. The LDGN is intricately structured, comprising stacked temporal linear layers and graph diffusion layers that harness the clarity of linear time series models. This configuration adeptly captures the nuanced interactions between graph-based diffusion processes and the dynamic spread of conditions across airport performances. Departing from conventional point forecasts, the LDGN produces a probabilistic output, prioritizing predictability and a strong capacity for generalization. The model’s pre-training is enhanced with stochastic mask reconstruction, a technique that significantly improves its ability to generalize. Through rigorous testing on real-world datasets, we have validated the LDGN’s superior performance in both long-term and very long-term forecasting. Our results demonstrate not only high accuracy and explainability but also a robust capacity for uncertainty quantification. Jing Yang 0017, Yi Lin 0006, Hongyu Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Spatiotemporal Propagation Learning for Network-Wide Flight Delay PredictionabstractAccurate and interpretable delay predictions are vital for decision-making in the aviation industry. However, effectively incorporating spatiotemporal dependencies and external factors related to delay propagation remains a challenge. To address this challenge, we propose the SpatioTemporal Propagation Network (STPN), a novel space-time separable graph convolutional network that models delay propagation by considering both spatial and temporal factors. STPN uses a multi-graph convolution model that considers both geographic proximity and airline schedules from a spatial perspective, while employing a multi-head self-attention mechanism that can be learned end-to-end and explicitly accounts for various types of temporal dependencies in delay time series from a temporal perspective. Experiments on two real-world delay datasets show that STPN outperforms state-of-the-art methods for multi-step ahead arrival and departure delay prediction in large-scale airport networks. Additionally, the counterfactuals generated by STPN provide evidence of its ability to learn explainable delay propagation patterns. Comprehensive experiments also demonstrate that STPN sets a robust benchmark for general spatiotemporal forecasting. The code for STPN is available athttps://github.com/Kaimaoge/STPN. Hongyu Yang 0002, Yi Lin 0006 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Unsupervised 3D Animal Canonical Pose Estimation with Geometric Self-SupervisionabstractAlthough analyzing animal shape and pose has potential applications in many fields, there is little work on 3D animal pose estimation. This can be attributed to two aspects: the lack of large-scale well-annotated datasets, and perspective ambiguities which make it difficult to map 2D space to 3D space. To address data scarcity, we propose an unsupervised method to estimate 3D animal pose, given only 2D poses. To deal with perspective ambiguities, we introduce a canonical consistency loss and a camera consistency loss to impose geometric priors in the training process, and combine the reprojection loss and the 2D pose discriminator to enable self-supervised learning. Specifically, given a 2D pose, the pose generator network generates a corresponding 3D pose and the camera network estimates a camera rotation. During training, the generated 3D pose is randomly reprojected onto camera viewpoints to synthesize a new 2D pose. The synthesized 2D pose is decomposed into a 3D pose and a camera rotation, based on which consistency losses are imposed in both 3D canonical poses and camera rotations for self-supervised training. We evaluate the proposed method on real and synthetic datasets, i.e., SMAL and AcinoSet. The experimental results demonstrate the effectiveness of the proposed method and we achieve state-of-the-art performance among unsupervised algorithms for 3D animal canonical pose estimation. Xiaowei Dai, Shuiwang Li, Qijun Zhao, Hongyu Yang 0002 |
FG | 4 |
| 2023 | GanNeXt: A New Convolutional GAN for Anomaly Detection
Bowei Pu, Shiyong Lan, Wenwu Wang 0001, Caiying Yang, Hongyu Yang 0002 |
ICANN (3) | 6 |
| 2023 | Visual-Haptic-Kinesthetic Object Recognition with Multimodal Transformer
Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Hongyu Yang 0002 |
ICANN (7) | 6 |
| 2023 | Flow-Based One-Class Anomaly Detection with Multi-Frequency Feature FusionabstractAnomaly detection in computer vision seeks to identify samples outside of a predefined distribution, including texture defect detection and semantic anomaly detection. However, existing methods are difficult to simultaneously achieve high performance for both types of anomaly detection. To address this issue, we propose a new flow-based anomaly detection method. Firstly, we use semantic features extracted from a pre-trained backbone to learn the distribution of normal data from a semantic perspective. Secondly, we introduce a multi-frequency feature fusion module to aggregate semantic and texture information, which substantially improves performance for both types of anomaly detection at the same time. Extensive experiments on multiple well-known datasets demonstrate that our proposed method performs well in both types of anomaly detection, specially, achieves state-of-the-art performance in one-class anomaly detection. The codes will be available at https://github.com/SYLan2019/FOAD-MFFF. Shiyong Lan, Weikang Huang, Yitong Ma, Hongyu Yang 0002, Yilin Zheng |
ICIP | 5 |
| 2023 | DLAHSD: Dynamic Label Adopted In Auxiliary Head for SAR DetectionabstractShip detection in synthetic aperture radar (SAR) images is a major issue in maritime surveillance and port management. Existing challenges are mainly as follows: (1) Tiny ships are mixed with scattered noise spots on the sea. (2) Ships are present in extreme aspect-ratios and various scales. (3) The land background blurs the outline of coastal ships. To address these problems, we propose an efficient detection neural network (DLAHSD) that integrates the Multi-scale Feature Location Fusion (MFLF) module and the Auxiliary Detection Head (ADH) based CenterNet. In addition, we designed a Dynamic Elliptic Gaussian (DEG) module to label the heatmap of ships. Experimental results on the challenging SSDD dataset show that our model offers improved performance over the baseline methods. The codes will be available at https://github.com/SYLan2019/DLAHSD. Xiaoxiao Yin, Shiyong Lan, Weikang Huang, Yitong Ma, Wenwu Wang 0001, Hongyu Yang 0002, Yilin Zheng |
ICIP | 6 |
| 2023 | A Semantics-Aware Normalizing Flow Model for Anomaly DetectionabstractAnomaly detection in computer vision aims to detect outliers from input image data. Examples include texture defect detection and semantic discrepancy detection. However, existing methods are limited in detecting both types of anomalies, especially for the latter. In this work, we propose a novel semantics-aware normalizing flow model to address the above challenges. First, we employ the semantic features extracted from a backbone network as the initial input of the normalizing flow model, which learns the mapping from the normal data to a normal distribution according to semantic attributes, thus enhances the discrimination of semantic anomaly detection. Second, we design a new feature fusion module in the normalizing flow model to integrate texture features and semantic features, which can substantially improve the fitting of the distribution function with input data, thus achieving improved performance for the detection of both types of anomalies. Extensive experiments on five well-known datasets for semantic anomaly detection show that the proposed method outperforms the state-of-the-art baselines. The codes will be available at https://github.com/SYLan2019/SANF-AD. Shiyong Lan, Weikang Huang, Wenwu Wang 0001, Hongyu Yang 0002, Yitong Ma, Yongjie Ma |
ICME | 5 |
| 2023 | Long Short-Term Perception Network for Dynamic Facial Expression Recognition
Chengcheng Lu, Yiben Jiang, Keren Fu, Qijun Zhao, Hongyu Yang 0002 |
PRCV (5) | 5 |
| 2023 | Differential evolution with variable leader-adjoint populations
Hongyu Yang 0002, Hu Chen 0002 |
Appl. Intell. | 3 |
| 2023 | Multiscale progressive text prompt network for medical image segmentation
Xianjun Han, Qianqian Chen 0007, Zhaoyang Xie, Xuejun Li 0001, Hongyu Yang 0002 |
Comput. Graph. | 5 |
| 2023 | Learning the degradation distribution for medical image superresolution via sparse swin transformer
Xianjun Han, Zhaoyang Xie, Qianqian Chen 0007, Xuejun Li 0001, Hongyu Yang 0002 |
Comput. Graph. | 5 |
| 2023 | Enhancing differential evolution algorithm using leader-adjoint populations
Hongyu Yang 0002, Hu Chen 0002, Bo Yang 0063 |
Inf. Sci. | 3 |
| 2022 | Learning Interval-Aware Embedding for Macro and Micro-expression Spotting
Wenchao Du, Hu Chen 0002, Hongyu Yang 0002 |
ACCV (4) | 5 |
| 2022 | Animal Pose Refinement in 2D Images with 3D Constraints
Xiaowei Dai, Shuiwang Li, Qijun Zhao, Hongyu Yang 0002 |
BMVC | 4 |
| 2022 | Face Super-Resolution with Spatial Attention Guided by Multiscale Receptive-Field Features
Weikang Huang, Shiyong Lan, Wenwu Wang 0001, Xuedong Yuan, Hongyu Yang 0002, Piaoyang Li |
ICANN (1) | 5 |
| 2022 | A Transformer-Based GAN for Anomaly Detection
Caiyin Yang, Shiyong Lan, Weikang Huang, Wenwu Wang 0001, Hongyu Yang 0002, Piaoyang Li |
ICANN (2) | 6 |
| 2022 | DSTAGNN: Dynamic Spatial-Temporal Aware Graph Neural Network for Traffic Flow ForecastingabstractAs a typical problem in time series analysis, traffic flow prediction is one of the most important application fields of machine learning. However, achieving highly accurate traffic flow prediction is a challenging task, due to the presence of complex dynamic spatial-temporal dependencies within a road network. This paper proposes a novel Dynamic Spatial-Temporal Aware Graph Neural Network (DSTAGNN) to model the complex spatial-temporal interaction in road network. First, considering the fact that historical data carries intrinsic dynamic information about the spatial structure of road networks, we propose a new dynamic spatial-temporal aware graph based on a data-driven strategy to replace the pre-defined static graph usually used in traditional graph convolution. Second, we design a novel graph neural network architecture, which can not only represent dynamic spatial relevance among nodes with an improved multi-head attention mechanism, but also acquire the wide range of dynamic temporal dependency from multi-receptive field features via multi-scale gated convolution. Extensive experiments on real-world data sets demonstrate that our proposed method significantly outperforms the state-of-the-art methods. Shiyong Lan, Yitong Ma, Weikang Huang, Wenwu Wang 0001, Hongyu Yang 0002, Pyang Li |
ICML | 5 |
| 2022 | Depth Completion Using Geometry-Aware EmbeddingabstractExploiting internal spatial geometric constraints of sparse LiDARs is beneficial to depth completion, however, has been not explored well. This paper proposes an efficient method to learn geometry-aware embedding, which encodes the local and global geometric structure information from 3D points, e.g., scene layout, object's sizes and shapes, to guide dense depth estimation. Specifically, we utilize the dynamic graph representation to model generalized geometric relationship from irregular point clouds in a flexible and efficient manner. Further, we joint this embedding and corresponded RGB appearance information to infer missing depths of the scene with well structure-preserved details. The key to our method is to integrate implicit 3D geometric representation into a 2D learning architecture, which leads to a better trade-off between the performance and efficiency. Extensive experiments demonstrate that the proposed method outperforms previous works and could reconstruct fine depths with crisp boundaries in regions that are over-smoothed by them. The ablation study gives more insights into our method that could achieve significant gains with a simple design, while having better generalization capability and stability. The code is available at https://github.com/Wenchao-Du/GAENet. Wenchao Du, Hu Chen 0002, Hongyu Yang 0002, Yi Zhang 0018 |
ICRA | 3 |
| 2022 | A backtracking differential evolution with multi-mutation strategies autonomy and collaboration
Bo Yang 0063, Hongyu Yang 0002, Miyi Zeng |
Appl. Intell. | 5 |
| 2022 | Ref-ZSSR: Zero-Shot Single Image Superresolution with Reference ImageabstractAbstract Single image superresolution (SISR) has achieved substantial progress based on deep learning. Many SISR methods acquire pairs of low‐resolution (LR) images from their corresponding high‐resolution (HR) counterparts. Being unsupervised, this kind of method also demands large‐scale training data. However, these paired images and a large amount of training data are difficult to obtain. Recently, several internal, learning‐based methods have been introduced to address this issue. Although requiring a large quantity of training data pairs is solved, the ability to improve the image resolution is limited if only the information of the LR image itself is applied. Therefore, we further expand this kind of approach by using similar HR reference images as prior knowledge to assist the single input image. In this paper, we proposed zero‐shot single image superresolution with a reference image (Ref‐ZSSR). First, we use an unconditional generative model to learn the internal distribution of the HR reference image. Second, a dual‐path architecture that contains a downsampler and an upsampler is introduced to learn the mapping between the input image and its downscaled image. Finally, we combine the reference image learning module and dual‐path architecture module to train a new generative model that can generate a superresolution (SR) image with the details of the HR reference image. Such a design encourages a simple and accurate way to transfer relevant textures from the reference high‐definition (HD) image to LR image. Compared with using only the image itself, the HD feature of the reference image improves the SR performance. In the experiment, we show that the proposed method outperforms previous image‐specific network and internal learning‐based methods. Xianjun Han, Xuejun Li 0001, Hongyu Yang 0002 |
Comput. Graph. Forum | 5 |
| 2022 | Discovering regression-detection bi-knowledge transfer for unsupervised cross-domain crowd counting
Yuting Liu 0004, Zheng Wang 0007, Miaojing Shi, Shin'ichi Satoh 0001, Qijun Zhao, Hongyu Yang 0002 |
Neurocomputing | 6 |
| 2022 | An Adaptive Specific Emitter Identification System for Dynamic Noise DomainabstractWith the sharply increasing number of wireless devices, specific emitter identification (SEI) technology based on radio frequency fingerprint (RFF) is employed to enhance devices’ security for authentication. However, the problems of poor performance in a highly dynamic interference environment, high dependence on quality of data sets, and less consideration of calculation resources are urgent to be resolved in a practical SEI applying domain. Our adaptive SEI system focuses on these problems, proposing a preprocessing algorithm improving synchrosqueezed wavelet transforms by energy regularization (ISWTE) to simplify the process of signal preprocessing while improving the robustness of RFF, an unsupervised neural network noise feature extracting GAN (NEGAN) to reduce the dependence of the data sets quality while obtaining precise clean RFF features from signals with noises and an optimized structure of waveform and classifier to achieve a better performance with less complexity and little cost. To test the system performance, we have generated a real wireless device data sets with ten devices, including a training set and a base test set both with SNR 10 dB, and a dynamic test set with SNR ranges −20 to 10 dB. The proposed adaptive SEI system by our work obtains a test accuracy of 0.96 at SNR 10 dB, 0.85 at SNR 0 dB, and 0.25 at SNR −20 dB only through SNR 10-dB training set. Most of the existing methods through neither of high SNR or enhanced mixed SNRs training set failed to complete devices identification under SNR −10 dB, and test accuracy of them deceased dramatically at SNR 0 dB. Miyi Zeng, Hongyu Yang 0002 |
IEEE Internet Things J. | 6 |
| 2021 | Small data assisting face image illumination normalization
Xianjun Han, Hongyu Yang 0002, Xuejun Li 0001 |
Comput. Graph. | 3 |
| 2021 | Long short-term memory self-adapting online random forests for evolving data stream regression
Hongyu Yang 0002, Yanci Zhang, Ping Li 0024, Cheng Ren |
Neurocomputing | 2 |
| 2021 | Online Rebuilding Regression Random Forests
Hongyu Yang 0002, Yanci Zhang, Ping Li 0024 |
Knowl. Based Syst. | 2 |
| 2020 | Learning Invariant Representation for Unsupervised Image RestorationabstractRecently, cross domain transfer has been applied for unsupervised image restoration tasks. However, directly applying existing frameworks would lead to domain-shift problems in translated images due to lack of effective supervision. Instead, we propose an unsupervised learning method that explicitly learns invariant presentation from noisy data and reconstructs clear observations. To do so, we introduce discrete disentangling representation and adversarial domain adaption into general domain transfer framework, aided by extra self-supervised modules including background and semantic consistency constraints, learning robust representation under dual domain constraints, such as feature and image domains. Experiments on synthetic and real noise removal tasks show the proposed method achieves comparable performance with other stateof-the-art supervised and unsupervised methods, while having faster and stable convergence than other domain adaption methods. Wenchao Du, Hu Chen 0002, Hongyu Yang 0002 |
CVPR | 3 |
| 2020 | Towards Unsupervised Crowd Counting via Regression-Detection Bi-knowledge TransferabstractUnsupervised crowd counting is a challenging yet not largely explored task. In this paper, we explore it in a transfer learning setting where we learn to detect and count persons in an unlabeled target set by transferring bi-knowledge learnt from regression- and detection-based models in a labeled source set. The dual source knowledge of the two models is heterogeneous and complementary as they capture different modalities of the crowd distribution. We formulate the mutual transformations between the outputs of regression- and detection-based models as two scene-agnostic transformers which enable knowledge distillation between the two models. Given the regression- and detection-based models and their mutual transformers learnt in the source, we introduce an iterative self-supervised learning scheme with regression-detection bi-knowledge transfer in the target. Extensive experiments on standard crowd counting benchmarks, ShanghaiTech, UCF_CC_50, and UCF_QNRF demonstrate a substantial improvement of our method over other state-of-the-arts in the transfer learning setting. Yuting Liu 0004, Zheng Wang 0007, Miaojing Shi, Shin'ichi Satoh 0001, Qijun Zhao, Hongyu Yang 0002 |
ACM Multimedia | 6 |
| 2020 | Normalization of face illumination with photorealistic texture via deep image prior synthesis
Xianjun Han, Yanli Liu 0002, Hongyu Yang 0002, Guanyu Xing, Yanci Zhang |
Neurocomputing | 3 |
| 2020 | Online random forests regression with memories
Hongyu Yang 0002, Yanci Zhang, Ping Li 0024 |
Knowl. Based Syst. | 2 |
| 2020 | Asymmetric Joint GANs for Normalizing Face Illumination From a Single ImageabstractIllumination normalization for face recognition is very important when a face is captured under harsh lighting conditions. Instead of designing hand-crafted features, in this paper we formulate face illumination normalization as an image-to-image translation task. A great challenge of face normalization is that human facial structures are particularly sensitive to image structure distortion, which frequently occurs in traditional image-to-image translation tasks. Unfortunately, sometimes even slight facial structure distortions may prohibit human eyes and machine face recognition methods from identifying face identities. To address this issue, a novel GAN- based network architecture called the asymmetric joint generative adversarial network (AJGAN) is developed to normalize face images under arbitrary illumination conditions, without known face geometry and albedo information. In addition, an illumination normalization GAN $G_1$ and an asymmetric relighting GAN $G_2$ that maps a frontal-illuminated image to images with various lighting conditions are incorporated in AJGAN to maintain personalized facial structures. To avoid image blurring caused by the under-constrained relighting mapping, we introduce a scheme of one-hot lighting labels into $G_2$ and enforce label classification loss. Furthermore, the number of training images starting from a very limited number of labels is dynamically extended by the combination of different lighting labels. Qualitative and quantitative experiments on three databases validate that AJGAN significantly outperforms the state-of-the-art methods. Xianjun Han, Hongyu Yang 0002, Guanyu Xing, Yanli Liu 0002 |
IEEE Trans. Multim. | 2 |
| 2019 | A differential evolution algorithm with dual preferred learning mutation
Meijun Duan, Hongyu Yang 0002 |
Appl. Intell. | 2 |
| 2019 | Visual Attention Network for Low-Dose CTabstractNoise and artifacts are intrinsic to low-dose computed tomography (LDCT) data acquisition, and will significantly affect the imaging performance. Perfect noise removal and image restoration is intractable in the context of LDCT due to the statistical and the technical uncertainties. In this letter, we apply the generative adversarial network (GAN) framework with a visual attention mechanism to deal with this problem in a data-driven/machine learning fashion. Our main idea is to inject visual attention knowledge into the learning process of GAN to provide a powerful prior of the noise distribution. By doing this, both the generator and discriminator networks are empowered with visual attention information so that they will not only pay special attention to noisy regions and surrounding structures but also explicitly assess the local consistency of the recovered regions. Our experiments qualitatively and quantitatively demonstrate the effectiveness of the proposed method with clinic CT images. Wenchao Du, Hu Chen 0002, Peixi Liao, Hongyu Yang 0002, Ge Wang 0001, Yi Zhang 0018 |
IEEE Signal Process. Lett. | 4 |
| 2018 | Self-adaptive differential evolution algorithm with improved mutation strategy
Hongyu Yang 0002 |
Soft Comput. | 3 |
| 2017 | Self-adaptive differential evolution algorithm with improved mutation mode
Hongyu Yang 0002 |
Appl. Intell. | 3 |
| 2016 | Efficient kd-tree construction for ray tracing using ray distribution sampling
Hongyu Yang 0002, Yanci Zhang |
Multim. Tools Appl. | 2 |
| 2015 | Evacuation Simulation Incorporating Safety Signs and Information Sharing
Yu Niu, Hongyu Yang 0002, Jianbo Fu, Xiaodong Che, Bin Shui, Yanci Zhang |
ICIG (2) | 2 |
| 2014 | Detecting soft shadows in a single outdoor image: From local edge-based models to global constraints
Yanli Liu 0002, Qijun Zhao, Hongyu Yang 0002 |
Comput. Graph. | 4 |