VLDB 2026 Research / reviewers in the wild / expert
Fengyuan Hu
dblp:96/11247
· DBLP profile ↗
17ranked-venue papers
6as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference under AmbiguitiesabstractSpatial expressions in situated communication can be ambiguous, as their meanings vary depending on the frames of reference (FoR) adopted by speakers and listeners. While spatial language understanding and reasoning by vision-language models (VLMs) have gained increasing attention, potential ambiguities in these models are still under-explored. To address this issue, we present the COnsistent Multilingual Frame Of Reference Test (COMFORT), an evaluation protocol to systematically assess the spatial reasoning capabilities of VLMs. We evaluate nine state-of-the-art VLMs using COMFORT. Despite showing some alignment with English conventions in resolving ambiguities, our experiments reveal significant shortcomings of VLMs: notably, the models (1) exhibit poor robustness and consistency, (2) lack the flexibility to accommodate multiple FoRs, and (3) fail to adhere to language-specific or culture-specific conventions in cross-lingual tests, as English tends to dominate other languages. With a growing effort to align vision-language models with human cognitive intuitions, we call for more attention to the ambiguous nature and cross-cultural diversity of spatial reasoning. Fengyuan Hu, Jayjun Lee, Freda Shi, Parisa Kordjamshidi, Joyce Y. Chai, Ziqiao Ma 0001 |
ICLR | 2 |
| 2025 | Deep learning-based fusion of nuclear segmentation features for microsatellite instability and tumor mutational burden prediction in digestive tract cancers: a multicenter validation studyabstractMicrosatellite instability (MSI) and tumor mutational burden (TMB) are crucial biomarkers in gastric (GC) and colorectal cancer (CRC), yet their conventional sequencing-based detection is costly and time-consuming. Since only ~20% of patients are MSI-high or TMB-high and likely to benefit from immunotherapy, expensive genomic testing is often unjustified. This study developed a deep learning framework to predict MSI and TMB status directly from routinely available Hematoxylin and Eosin (H&E)-stained whole-slide images, leveraging fused nuclear segmentation features to improve accuracy. Using samples from TCGA (350 GC and 376 CRC for MSI; 400 GC and 387 CRC for TMB), image features were extracted with CLAM and nuclear features with Hover-Net. These features were combined via Multimodal Compact Bilinear Pooling and utilized in six distinct deep learning models. By fusing the nucleus segmentation features, the model increased area under the receiver operating characteristic curve (AUC) by 1%-3% and recall by 5%-11% in five-fold cross-validation, significantly outperforming models that relied solely on image features. External validation on a CRC dataset from the China-Japan Friendship hospital further validated the model's robustness, achieving an AUC of 0.81 and a recall of 0.80 for MSI prediction. Additionally, notable differences in cellular composition were observed across cancer types and clinical groups, emphasizing the pivotal role of cellular features in cancer development. These findings highlight the advantages of integrating H&E-stained image features with nuclear segmentation data and advanced deep learning techniques to improve predictive accuracy and reduce the cost of MSI/TMB testing, potentially advancing personalized cancer treatment strategies. Jiaying Han, Fengyuan Hu, Geng Tian, Dingrong Zhong, Jialiang Yang |
Briefings Bioinform. | 4 |
| 2025 | Keypoint-Based SAR Structure From Motion via Riemannian OptimizationabstractStructure-from-motion (SfM) is the concept of estimating both sensor pose and three-dimensional (3D) scene structure from input images. The difficulty with synthetic aperture radar (SAR) SfM stems from the non-linearity of radar imaging, making it hard to decouple and calculate radar pose and structure. Existing methods typically address this issue by simplifying the imaging process or introducing auxiliary data. In this paper, we propose to jointly solve radar pose and structure by formulating an optimization problem, which only needs two-dimensional (2D) SAR observations as input. The objective function is written as the summation of all reprojection errors established by the precise range-Doppler (RD) imaging model, with radar poses embedded into the transformations between the sensor and scene coordinate systems. The rotational components in radar poses are represented as matrices and constrained on the special orthogonal groupSO(3). A special orthogonal Riemannian conjugate gradient algorithm (SO-RCG) is then proposed to solve the optimization problem. The proposed algorithm preserves the orthogonality of rotation matrices and updates all variables iteratively. Furthermore, we deduce that only five degrees of freedom (DoFs) in radar pose are involved in determining the imaging results of targets. We also discuss the ambiguity issue in multiview SAR observation leading to local minimums and introduce strategies against ambiguity. Experimental results on real-measured and simulated datasets show that the proposed algorithm is effective on both near- and far-field cases, and can be applied to both side-looking and squint SAR. Fengyuan Hu, Xue Jiang 0001, Junfeng Wang 0001, Xingzhao Liu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional PropertiesabstractA major reason behind the recent success of large language models (LLMs) is their incontext learning capability, which makes it possible to rapidly adapt them to downstream textbased tasks by prompting them with a small number of relevant demonstrations.While large vision-language models (VLMs) have recently been developed for tasks requiring both text and images, they largely lack in-context learning over visual information, especially in understanding and generating text about videos.In this work, we implement Emergent In-context Learning on Videos (EILeV), a novel training paradigm that induces in-context learning over video and text by capturing key properties of pre-training data found by prior work to be essential for in-context learning in transformers.In our experiments, we show that EILeV-trained models outperform other off-the-shelf VLMs in few-shot video narration for novel, rare actions.Furthermore, we demonstrate that these key properties of bursty distributions, skewed marginal distributions, and dynamic meaning each contribute to varying degrees to VLMs' in-context learning capability in narrating procedural videos.Our results, analysis, and EILeV-trained models yield numerous insights about the emergence of in-context learning over video and text, creating a foundation for future work to optimize and scale VLMs for open-domain video understanding and reasoning.1 Keunwoo Peter Yu, Fengyuan Hu, Shane Storks, Joyce Y. Chai |
EMNLP | 3 |
| 2024 | SAR Pose Estimation with Circular-N-Point: A Two-Step MethodabstractThis paper proposes a fast two-step method which aims to address the SAR pose estimation problem and enable Unmanned Aerial Vehicles (UAVs) the self-localization capability under harsh conditions. Firstly, the monocular SAR pose estimation is formulated as a Circular-n-Point (CnP) problem based on frequency-domain SAR imaging mechanism. Then, we propose to decouple the motion components of radar platform and solve the overdetermined equations of direction and location sequentially. Experimental results validate the effectiveness, accuracy, and efficiency of the proposed method. Fengyuan Hu, Xue Jiang 0001, Junfeng Wang 0001, Xingzhao Liu, Lingyu Wang 0004 |
IGARSS | 1 |
| 2024 | A Transformer-Based Optronic Neural Network for SAR Target RecognitionabstractTransformer has shown great capability in remote sensing and automatic target recognition (ATR). Due to the self-attention mechanism, the Transformer could extract global features while parallelizing training. However, the computational costs and power consumption are challenging the electronic computing techniques. Here, we develop a Transformer-based optronic neural network (TOPNN) for synthetic aperture radar (SAR) target recognition. We implement the self-attention mechanism in optics, significantly reducing the network computational costs. Compared with digital techniques, the TOPNN promises the speed of light, low computational costs, and low power consumption. Experiments on the Moving and Stationary Target Acquisition and Recognition (MSTAR) dataset demonstrate the feasibility and efficiency of TOPNN for SAR target recognition. Fengyuan Hu, Jiahui Ma, Yesheng Gao, Xingzhao Liu |
IGARSS | 2 |
| 2024 | Speckle-Based Residual Optronic Convolutional Neural Network for SAR Target Recognition in Scattering Imaging ScenariosabstractScattering imaging is a pervasive scenario in many areas, especially challenging the performance of remote sensing and automatic target recognition (ATR). Recently, deep learning was utilized for synthetic aperture radar (SAR) ATR in scattering scenarios by extracting the feature of speckle patterns. However, huge computational costs and power consumption challenge its development. Here, we develop a speckle-based residual optronic convolutional neural network (S-ROPCNN) for SAR target recognition. Specifically, we model the light scattering scenarios and build the optical imaging system to produce the speckle patterns for network training. The S-ROPCNN performs SAR target recognition in optical platforms with the speed of light, low computational cost, and low energy consumption. Experiments on the Moving and Stationary Target Acquisition and Recognition (MSTAR) dataset demonstrate the feasibility of S-ROPCNN for SAR target recognition in scattering imaging scenarios. Fengyuan Hu, Guozheng Xu, Mengyang Shi, Yesheng Gao |
IGARSS | 2 |
| 2024 | A Linearly Continous False Alarm Removing Method in a Multichannel SAR-GMTI SystemabstractIn this paper, a method aimed at removing the linearly continuous false alarm caused by the artificial building with unignored elevation in a multichannel synthetic aperture radar (SAR) system with ground moving target indication (GMTI) mode is proposed. In the proposed method, the Hough transform is firstly utilized to obtain the potential false alarms during the iteration procedure. After that, the linearly continuous false alarms are finally obtained and removed based on the phase similarity characteristic of those building scatters. Real-measured SAR data processing results are presented to verity the effectiveness and feasibility of the proposed method. Lingyu Wang 0004, Xin Li 0005, Penghui Huang, Haojuan Yuan, Changhong He, Muyang Zhan, Fengyuan Hu |
IGARSS | 9 |
| 2024 | STAP Performance Analysis Using Different Receiving Antenna Weighting Manners in Space-Based Early Warning Radar SystemsabstractIn this paper, the influence of different receiving array element weighting manners on the space-time adaptive processing (STAP) technique is analyzed. Firstly, the element-based echo signal in a space-borne early warning radar (SBEWR) system is established, and then the STAP procedure by utilizing two alternative receiving array element weighting models are introduced. After that, the antenna patterns with respect to two weighting manners are discussed and the target output signal-to-clutter-plus-noise ratio (SCNR) after STAP are evaluated. The numerical results show that the subarray element weighting manner obtains a better STAP performance without channel mismatch compensation while a slightly inferior STAP performance with channel mismatch compensation compared with the whole antenna weighting manner. Lingyu Wang 0004, Xin Lin 0002, Haojuan Yuan, Penghui Huang, Changhong He, Muyang Zhan, Fengyuan Hu |
IGARSS | 8 |
| 2024 | DGA: Direction-Guided Attack Against Optical Aerial Detection in Camera Shooting Direction-Agnostic ScenariosabstractPatch-based adversarial attacks have increasingly aroused concerns due to their application potential in military and civilian fields. In aerial imagery, numerous targets exhibit inherent directionality, such as vehicles and ships, giving rise to the emergence of oriented object detection tasks; similarly, adversarial patches also exhibit intrinsic orientation due to their lack of perfect symmetry. Existing methods presuppose a static alignment between the adversarial patch’s orientation and the camera’s coordinate system – an assumption that is frequently violated in aerial images, whose effectiveness degrades in real-world scenarios. In this paper, we investigate the often-neglected aspect of patch orientation in adversarial attacks and its impact on camouflage effectiveness, particularly when the orientation is not congruent with the target. A new Directional Guided Attack (DGA) framework is proposed for deceiving real-world aerial detectors, which shows robust and adaptable attack performance in camera shooting direction agnostic (CSDA) scenarios. The core idea of DGA is to utilize affine transformations to constrain the relative orientation of the patch to the target and introduce three types of loss to reduce target detection confidence, make the color printable, and smooth the patch color. We introduce a direction-guided evaluation methodology to bridge the gap between patch performance in the digital domain and its actual real-world efficacy. Moreover, we establish a drone-based vehicle detection dataset (SJTU-4K), which labels the orientation of the target, to assess the robustness of patches under various shooting altitudes and views. Extensive proportionally scaled and 1:1 experiments are performed in physical scenarios, demonstrating the superiority and potential of the proposed framework for real-world attacks. Yue Zhou 0005, Shu-Qi Sun, Xue Jiang 0001, Guozheng Xu, Fengyuan Hu, Xingzhao Liu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense ReasoningabstractPre-trained language models (PLMs) have shown impressive performance in various language tasks.However, they are prone to spurious correlations, and often generate illusory information.In real-world applications, PLMs should justify decisions with formalized, coherent reasoning chains, but this challenge remains under-explored.Cognitive psychology theorizes that humans are capable of utilizing fast and intuitive heuristic thinking to make decisions based on past experience, then rationalizing the decisions through slower and deliberative analytic reasoning.We incorporate these interlinked dual processes in fine-tuning and in-context learning with PLMs, applying them to two language understanding tasks that require coherent physical commonsense reasoning.We show that our proposed Heuristic-Analytic Reasoning (HAR) strategies drastically improve the coherence of rationalizations for model decisions, yielding state-of-the-art results on Tiered Reasoning for Intuitive Physics (TRIP).We also find that this improved coherence is a direct result of more faithful attention to relevant language context in each step of reasoning.Our findings suggest that human-like reasoning strategies can effectively improve the coherence and reliability of PLM reasoning. Shane Storks, Fengyuan Hu, Sungryull Sohn, Moontae Lee, Honglak Lee, Joyce Y. Chai |
EMNLP | 3 |
| 2023 | SAR Structure-From-Motion via Matrix FactorizationabstractStructure-from-Motion (SfM) is the process of estimating 3D scene structure and sensor pose from a set of 2D inputs. In this paper, we present a matrix factorization scheme for solving SAR SfM problem. First, SAR imaging model is linearized at local scene and the SfM problem is converted to a problem that decomposes data matrices into the product of two kinds of matrices, one of which satisfies Stiefel constraint. We then propose a Riemannian conjugate gradient descent algorithm leveraging the Stiefel constraint. SAR SfM is solved by the proposed algorithm via alternating iteratively estimating radar pose and scene structure. Finally, the feasibility and accuracy of the proposed scheme are verified through experiments. Fengyuan Hu, Xue Jiang 0001, Junfeng Wang 0001, Xingzhao Liu |
IGARSS | 1 |
| 2022 | Radar Pose Estimation and Structure-from-Motion for Airborne Circular VideoSARabstractA working video synthetic aperture radar (VideoSAR) can obtain continuous observations of a region of interest from different viewpoints, meaning those observations naturally contain 3D information for positioning. This paper reports our preliminary studies on SAR scene structure-from-motion using several frames extracted from a VideoSAR sequence. In contrast to classic stereo-radargrammetric workflow, the radar pose is estimated only with information measured by radar. The coordinates of each scattering point are calculated by least-square optimization. We also develop an affine transformation aware dense matching framework to accurately measure pixel-level correspondence between frames. Experiments on real VideoSAR data demonstrate the validity of our approach. Fengyuan Hu, Xue Jiang 0001, Junfeng Wang 0001, Xingzhao Liu |
IGARSS | 1 |
| 2022 | Correction to: ORFLine: a bioinformatic pipeline to prioritize small open reading frames identifies candidate secreted small proteins from lymphocytesabstractBioinformatics, 2021, 37(19), 3152–3159. https://doi.org/10.1093/bioinformatics/btab339 In the originally published version of this manuscript, the Associate Editor’s name, Jinbo Xu, was listed as an author in error. The publisher apologizes for the error. This error has been corrected online. Fengyuan Hu, Louise S. Matheson, Manuel D. Díaz-Muñoz, Alexander Saveliev, Martin Turner |
Bioinform. | 1 |
| 2021 | ORFLine: a bioinformatic pipeline to prioritize small open reading frames identifies candidate secreted small proteins from lymphocytesabstractMOTIVATION: The annotation of small open reading frames (smORFs) of <100 codons (<300 nucleotides) is challenging due to the large number of such sequences in the genome. RESULTS: In this study, we developed a computational pipeline, which we have named ORFLine, that stringently identifies smORFs and classifies them according to their position within transcripts. We identified a total of 5744 unique smORFs in datasets from mouse B and T lymphocytes and systematically characterized them using ORFLine. We further searched smORFs for the presence of a signal peptide, which predicted known secreted chemokines as well as novel micropeptides. Four novel micropeptides show evidence of secretion and are therefore candidate mediators of immunoregulatory functions. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at https://github.com/boboppie/ORFLine. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fengyuan Hu, Louise S. Matheson, Manuel D. Díaz-Muñoz, Alexander Saveliev, Jinbo Xu, Martin Turner |
Bioinform. | 1 |
| 2015 | A comparative study of RNA-seq analysis strategiesabstractThree principal approaches have been proposed for inferring the set of transcripts expressed in RNA samples using RNA-seq. The simplest approach uses curated annotations, which assumes the transcripts in a sample are a subset of the transcripts listed in a curated database. A more ambitious method involves aligning reads to a reference genome and using the alignments to infer the transcript structures, possibly with the aid of a curated transcript database. The most challenging approach is to assemble reads into putative transcripts de novo without the aid of reference data. We have systematically assessed the properties of these three approaches through a simulation study. We have found that the sensitivity of computational transcript set estimation is severely limited. Computational approaches (both genome-guided and de novo assembly) produce a large number of artefacts, which are assigned large expression estimates and absorb a substantial proportion of the signal when performing expression analysis. The approach using curated annotations shows good expression correlation even when the annotations are incomplete. Furthermore, any incorrect transcripts present in a curated set do not absorb much signal, so it is preferable to have a curation set with high sensitivity than high precision. Software to simulate transcript sets, expression values and sequence reads under a wider range of parameter values and to compare sensitivity, precision and signal-to-noise ratios of different methods is freely available online (https://github.com/boboppie/RSSS) and can be expanded by interested parties to include methods other than the exemplars presented in this article. Jürgen Jänes, Fengyuan Hu, Alex Lewin, Ernest Turro |
Briefings Bioinform. | 2 |
| 2012 | InterMine: a flexible data warehouse system for the integration and analysis of heterogeneous biological dataabstractSUMMARY: InterMine is an open-source data warehouse system that facilitates the building of databases with complex data integration requirements and a need for a fast customizable query facility. Using InterMine, large biological databases can be created from a range of heterogeneous data sources, and the extensible data model allows for easy integration of new data types. The analysis tools include a flexible query builder, genomic region search and a library of 'widgets' performing various statistical analyses. The results can be exported in many commonly used formats. InterMine is a fully extensible framework where developers can add new tools and functionality. Additionally, there is a comprehensive set of web services, for which client libraries are provided in five commonly used programming languages. AVAILABILITY: Freely available from http://www.intermine.org under the LGPL license. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Richard N. Smith, Jelena Aleksic, Daniela Butano, Adrian Carr, Sergio Contrino, Fengyuan Hu, Mike Lyne, Rachel Lyne, Alex Kalderimis, Kim Rutherford, Radek Stepan, Julie M. Sullivan, Matthew Wakeling, Xavier Watkins, Gos Micklem |
Bioinform. | 6 |