Fanjie Kong

dblp:197/2743 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Beyond Speaker Identity: Text Guided Target Speech Extraction
abstract
Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker’s identity like enrollment audio, face images, or videos, which may not always be available. In this paper, we propose a text-guided TSE model StyleTSE that uses natural language descriptions of speaking style in addition to the audio clue to extract the desired speech from a given mixture. Our model integrates a speech separation network adapted from SepFormer with a bi-modality clue network that flexibly processes both audio and text clues. To train and evaluate our model, we introduce a new dataset TextrolMix with speech mixtures and natural language descriptions. Experimental results demonstrate that our method effectively separates speech based not only on who is speaking, but also on how they are speaking, enhancing TSE when traditional audio clues are absent. Demos are at: https://mingyue66.github.io/TextrolMix/demo/
Mingyue Huo, Cong Phuoc Huynh, Fanjie Kong, Pichao Wang, Vimal Bhat
ICASSP4
2025 VLR-Driver: Large Vision-Language-Reasoning Models for Embodied Autonomous Driving
Fanjie Kong, Weihuang Chen, Zhongyu Guo, Hongbin Sun 0001
ICCV1
2025 Detect, Disambiguate, and Translate: On-Demand Visual Reasoning for Multimodal Machine Translation with Large Vision-Language Models
abstract
Danyang Liu, Fanjie Kong, Xiaohang Sun, Dhruva Patil, Avijit Vajpayee, Zhu Liu, Vimal Bhat, Najmeh Sadoughi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Fanjie Kong, Xiaohang Sun, Dhruva Patil, Avijit Vajpayee, Vimal Bhat, Najmeh Sadoughi
NAACL (Long Papers)2
2024 Hyperbolic Learning with Synthetic Captions for Open-World Detection
abstract
Open-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption datasets for training, which are extremely expensive to collect. Instead, we propose to transfer knowledge from vision-language models (VLMs) to enrich the open-vocabulary descriptions au-tomatically. Specifically, we bootstrap dense synthetic captions using pretrained VLMs to provide rich descriptions on different regions in images, and incorporate these captions to train a novel detector that generalizes to novel concepts. To mitigate the noise caused by hallucination in syn-thetic captions, we also propose a novel hyperbolic vision-language learning approach to impose a hierarchy between visual and caption embeddings. We call our detector “Hy-perLearner”. We conduct extensive experiments on a wide variety of open-world detection benchmarks (COCO, LVIS, Object Detection in the Wild, RefCoCo) and our results show that our model consistently outperforms existing state-of-the-art methods, such as GLIP, GLIPv2 and Grounding DINO, when using the same backbone.
Fanjie Kong, Yanbei Chen, Jiarui Cai, Davide Modolo
CVPR1
2023 Neural Insights for Digital Marketing Content Design
abstract
In digital marketing, experimenting with new website content is one of the key levers to improve customer engagement. However, creating successful marketing content is a manual and time-consuming process that lacks clear guiding principles. This paper seeks to close the loop between content creation and online experimentation by offering marketers AI-driven actionable insights based on historical data to improve their creative process. We present a neural-network-based system that scores and extracts insights from a marketing content design. Namely, a multimodal neural network predicts the attractiveness of marketing contents, and a post-hoc attribution method generates actionable insights for marketers to improve their content in specific marketing locations. Our insights not only point out the advantages and drawbacks of a given current content, but also provide design recommendations based on historical data. We show that our scoring model and insights work well both quantitatively and qualitatively.
Fanjie Kong, Yuan Li 0031, Houssam Nassif, Tanner Fiez, Ricardo Henao, Shreya Chakrabarti
KDD1
2023 Mitigating Test-Time Bias for Fair Image Retrieval
abstract
We address the challenge of generating fair and unbiased image retrieval results given neutral textual queries (with no explicit gender or race connotations), while maintaining the utility (performance) of the underlying vision-language (VL) model. Previous methods aim to disentangle learned representations of images and text queries from gender and racial characteristics. However, we show these are inadequate at alleviating bias for the desired equal representation result, as there usually exists test-time bias in the target retrieval set. So motivated, we introduce a straightforward technique, Post-hoc Bias Mitigation (PBM), that post-processes the outputs from the pre-trained vision-language model. We evaluate our algorithm on real-world image search datasets, Occupation 1 and 2, as well as two large-scale image-text datasets, MS-COCO and Flickr30k. Our approach achieves the lowest bias, compared with various existing bias-mitigation methods, in text-based image retrieval result while maintaining satisfactory retrieval performance. The source code is publicly available at \url{https://github.com/timqqt/Fair_Text_based_Image_Retrieval}.
Fanjie Kong, Weituo Hao, Ricardo Henao
NeurIPS1
2022 Efficient Classification of Very Large Images with Tiny Objects
abstract
An increasing number of applications in computer vision, specially, in medical imaging and remote sensing, become challenging when the goal is to classify very large images with tiny informative objects. Specifically, these classification tasks face two key challenges: i) the size of the input image is usually in the order of mega- or giga-pixels, however, existing deep architectures do not easily operate on such big images due to memory constraints, consequently, we seek a memory-efficient method to process these images; and ii) only a very small fraction of the input images are informative of the label of interest, resulting in low region of interest (ROI) to image ratio. However, most of the current convolutional neural networks (CNNs) are designed for image classification datasets that have relatively large ROIs and small image sizes (sub-megapixel). Existing approaches have addressed these two challenges in isolation. We present an end-to-end CNN model termed Zoom-In network that leverages hierarchical attention sampling for classification of large images with tiny objects using a single GPU. We evaluate our method on four large-image histopathology, road-scene and satellite imaging datasets, and one gigapixel pathology dataset. Experimental results show that our model achieves higher accuracy than existing methods while requiring less memory resources.
Fanjie Kong, Ricardo Henao
CVPR1
2021 Physics-Enhanced Machine Learning for Virtual Fluorescence Microscopy
abstract
This paper introduces a new method of data-driven microscope design for virtual fluorescence microscopy. We use a deep neural network (DNN) to effectively design optical patterns for specimen illumination that substantially improve upon the ability to infer fluorescence image information from unstained microscope images. To achieve this design, we include an illumination model within the DNN’s first layers that is jointly optimized during network training. We validated our method on two different experimental setups, with different magnifications and sample types, to show a consistent improvement in performance as compared to conventional microscope imaging methods. Additionally, to understand the importance of learned illumination on the inference task, we varied the number of illumination patterns being optimized (and thus the number of unique images captured) and analyzed how the structure of the patterns changed as their number increased. This work demonstrates the power of programmable optical elements at enabling better machine learning algorithm performance and at providing physical insight into next generation of machine-controlled imaging systems.
Colin L. V. Cooke, Fanjie Kong, Amey Chaware, Kevin C. Zhou, Kanghyun Kim, D. Michael Ando, Samuel J. Yang, Pavan Chandra Konda, Roarke Horstmeyer
ICCV2
2020 The Synthinel-1 dataset: a collection of high resolution synthetic overhead imagery for building segmentation
abstract
Recently deep learning - namely convolutional neural networks (CNNs) - have yielded impressive performance for the task of building segmentation on large overhead (e.g., satellite) imagery benchmarks. However, these benchmark datasets only capture a small fraction of the variability present in real-world overhead imagery, limiting the ability to properly train, or evaluate, models for real-world application. Unfortunately, developing a dataset that captures even a small fraction of real-world variability is typically infeasible due to the cost of imagery, and manual pixel-wise labeling of the imagery. In this work we develop an approach to rapidly and cheaply generate large and diverse synthetic overhead imagery for training segmentation CNNs. Using this approach, we generate and publicly-release a collection of synthetic overhead imagery, termed Synthinel-1, with full pixel-wise building labels. We use several benchmark datasets to demonstrate that Synthinel-1 is consistently beneficial when used to augment real-world training imagery, especially when CNNs are tested on novel geographic locations or conditions.
Fanjie Kong, Bohao Huang, Kyle Bradbury, Jordan M. Malof
WACV1
2019 Training a single multi-class convolutional segmentation network using multiple datasets with heterogeneous labels: preliminary results
abstract
Segmentation convolutional neural networks (CNNs) are now popular for the semantic segmentation (i.e., dense pixel-wise labeling) of remote sensing imagery, such as color or hyperspectral satellite imagery. In recent years a large number of hand-labeled datasets of overhead imagery have emerged, leading to breakthrough performance for CNNs. However, these datasets are typically used in isolation of one another because they are either (i) annotated with heterogeneous object type labels, or (ii) they are collected over different geographic areas. This imposes a major bottleneck on the value of these datasets. In this work we present what we call a class-asymmetric loss function that makes it possible to train a single multi-class network using multiple datasets that are heterogeneously-labeled. We show, for example, that it is possible to train a segmentation algorithm for Buildings, roads, and background using two datasets: one annotated with buildings and one annotated with buildings. We propose a class asymmetric loss that under certain common conditions, allows for one to train models on datasets in which the target class is unlabeled.
Fanjie Kong, Bohao Huang, Leslie M. Collins, Kyle Bradbury, Jordan M. Malof
IGARSS1
2019 Restoration algorithm for noisy complex illumination
abstract
Although promising results have been achieved in the restoration of complex illumination images with the Retinex algorithm, there are still some drawbacks in the processing of Retinex. Considering the noise characteristics of complex illumination images, in this study, we propose a novel restoration algorithm for noisy complex illumination, which combines guided adaptive multi‐scale Retinex (GAMSR) and improvement BayesShrink threshold filtering (IBTF) based on double‐density dual‐tree complex wavelet transform (DDDTCWT) domain. Extensive restoration experiments are conducted on three typical types images and the same image with different noises. On the basis of a series of evaluation indexes, we compare our method to those of state‐of‐the‐art algorithms. The results show that (i) SSIM of the proposed IBTF is superior to traditional Bayes threshold method by 15% as the standard variance is 100. (ii) PSNR of the proposed GAMSR enhances 15% to traditional MSR. (iii) The clarity of final results for restoration speeds up three times than that of original images, and the information entropy is improved slightly too. Therefore, the proposed method can effectively enhance the details, edges and textures of the image under complex illumination and noises.
Zhanwen Liu, Tao Gao 0001, Fanjie Kong, Ziheng Jiao, Aodong Yang, Bo Liu 0006
IET Comput. Vis.3