EDBT 2026 Demo / reviewers in the wild / expert
Fei Jia
dblp:121/0045
· DBLP profile ↗
20ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A self-supervised learning-based method for tunnel leakage defect few-shot recognition and visual explanationabstractThe frequently used supervised learning (SL) methods for detecting tunnel leakage defect require a large amount of labeled data, which is costly and challenging for practical applications. Moreover, benchmarking model performance faces challenge of the inconsistent annotation standards of dataset. To address these limitations, based on the self-established tunnel leakage instance segmentation dataset, this study integrates self-supervised learning with YOLO11n model to facilitate few-shot leakage recognition. In the workflow, YOLO-SimSiam is constructed and pre-trained using unlabeled images to enhance leakage representation learning, and a comprehensive evaluation method is proposed to assess feature extraction performance. Subsequently, YOLO11n model with pre-trained weights are fine-tuned on a limited number of labeled images for precise leakage recognition through transfer learning. Experiment results show that the model achieves competitive performance with only 10% dataset (0.910) comparable to that of SL on entire dataset (0.930). Besides, a theoretical conversion formula between leakage pixel number and real area is deduced and verified in a field calibration experiment with an error rate of 1.759%, providing a normative defect severity evaluation measure. The proposed method can effectively reduce the dependency on extensive data annotation, and thus promote the application of deep learning in practical tunnel maintenance. Fei Jia, Ya-Dong Xue, Yu-xuan Li, Yong-Fa Guo |
Adv. Eng. Informatics | 1 |
| 2025 | SWAN: An Efficient and Scalable Approach for Long-Context Language ModelingabstractKrishna C Puvvada, Faisal Ladhak, Santiago Akle Serano, Cheng-Ping Hsieh, Shantanu Acharya, Somshubra Majumdar, Fei Jia, Samuel Kriman, Simeng Sun, Dima Rekesh, Boris Ginsburg. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Krishna C. Puvvada, Faisal Ladhak, Santiago Akle Serano, Cheng-Ping Hsieh, Shantanu Acharya, Somshubra Majumdar, Fei Jia, Samuel Kriman, Simeng Sun, Dima Rekesh, Boris Ginsburg |
EMNLP | 7 |
| 2025 | Star Attention: Efficient LLM Inference over Long SequencesabstractInference with Transformer-based Large Language Models (LLMs) on long sequences is both costly and slow due to the quadratic complexity of the self-attention mechanism. We introduce Star Attention, a two-phase block-sparse approximation that improves computational efficiency by sharding attention across multiple hosts while minimizing communication overhead. In the first phase, the context is processed using blockwise-local attention across hosts, in parallel. In the second phase, query and response tokens attend to all prior cached tokens through sequence-global attention. Star Attention integrates seamlessly with most Transformer-based LLMs trained with global attention, reducing memory requirements and inference time by up to 11x while preserving 97-100% of accuracy. Shantanu Acharya, Fei Jia, Boris Ginsburg |
ICML | 2 |
| 2024 | Transducers with Pronunciation-Aware Embeddings for Automatic Speech RecognitionabstractThis paper proposes Transducers with Pronunciation-aware Embeddings (PET). Unlike conventional Transducers where the decoder embeddings for different tokens are trained independently, the PET model’s decoder embedding incorporates shared components for text tokens with the same or similar pronunciations. With experiments conducted in multiple datasets in Mandarin Chinese and Korean, we show that PET models consistently improve speech recognition accuracy compared to conventional Transducers. Our investigation also uncovers a phenomenon that we call error chain reactions. Instead of recognition errors being evenly spread throughout an utterance, they tend to group together, with subsequent errors often following earlier ones. Our analysis shows that PET models effectively mitigate this issue by substantially reducing the likelihood of the model generating additional errors following a prior one. Our implementation will be open-sourced with the NeMo toolkit. Hainan Xu, Zhehuai Chen, Fei Jia, Boris Ginsburg |
ICASSP | 3 |
| 2024 | OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning DatasetabstractRecent work has shown the immense potential of synthetically generated datasets for training large language models (LLMs), especially for acquiring targeted skills. Current large-scale math instruction tuning datasets such as MetaMathQA (Yu et al., 2024) and MAmmoTH (Yue et al., 2024) are constructed using outputs from closed-source LLMs with commercially restrictive licenses. A key reason limiting the use of open-source LLMs in these data generation pipelines has been the wide gap between the mathematical skills of the best closed-source LLMs, such as GPT-4, and the best open-source LLMs. Building on the recent progress in open-source LLMs, our proposed prompting novelty, and some brute-force scaling, we construct OpenMathInstruct-1, a math instruction tuning dataset with 1.8M problem-solution pairs. The dataset is constructed by synthesizing code-interpreter solutions for GSM8K and MATH, two popular math reasoning benchmarks, using the recently released and permissively licensed Mixtral model. Our best model, OpenMath-CodeLlama-70B, trained on a subset of OpenMathInstruct-1, achieves a score of 84.6% on GSM8K and 50.7% on MATH, which is competitive with the best gpt-distilled models. We will release our code, models, and the OpenMathInstruct-1 dataset under a commercially permissive license. Shubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, Igor Gitman |
NeurIPS | 5 |
| 2024 | Romanization Encoding For Multilingual ASRabstractWe introduce romanization encoding for script-heavy languages to optimize multilingual and code-switching Automatic Speech Recognition (ASR) systems. By adopting romanization encoding alongside a balanced concatenated tokenizer within a FastConformer-RNNT framework equipped with a Roman2Char module, we significantly reduce vocabulary and output dimensions, enabling larger training batches and reduced memory consumption. Our method decouples acoustic modeling and language modeling, enhancing the flexibility and adaptability of the system. In our study, applying this method to Mandarin-English ASR resulted in a remarkable 63.51% vocabulary reduction and notable performance gains of 13.72% and 15.03% on SEAME code-switching benchmarks. Ablation studies on MandarinKorean and Mandarin-Japanese highlight our method’s strong capability to address the complexities of other script-heavy languages, paving the way for more versatile and effective multilingual ASR systems. Wen Ding 0005, Fei Jia, Hainan Xu, Yu Xi, Junjie Lai, Boris Ginsburg |
SLT | 2 |
| 2023 | Accidental Learners: Spoken Language Identification in Multilingual Self-Supervised ModelsabstractIn this paper, we extend previous self-supervised approaches for language identification by experimenting with Conformer based architecture in a multilingual pre-training paradigm. We find that pre-trained speech models optimally encode language discriminatory information in lower layers. Further, we demonstrate that the embeddings obtained from these layers are significantly robust to classify unseen languages and different acoustic environments without additional training. After fine-tuning a pre-trained Conformer model on the VoxLin-gua107 dataset, we achieve results similar to current state-of-the-art systems for language identification. More, our model accomplishes this with 5x less parameters. We open-source the model through the NVIDIA NeMo toolkit. Travis M. Bartley, Fei Jia, Krishna C. Puvvada, Samuel Kriman, Boris Ginsburg |
ICASSP | 2 |
| 2023 | Multi-Blank Transducers for Speech RecognitionabstractThis paper proposes a modification to RNN-Transducer (RNN-T) models for automatic speech recognition (ASR). In standard RNN-T, the emission of a blank symbol consumes exactly one input frame; in our proposed method, we introduce additional blank symbols, which consume two or more input frames when emitted. We refer to the added symbols as big blanks, and the method multi-blank RNN-T. For training multi-blank RNN-Ts, we propose a novel logit under-normalization method in order to prioritize emissions of big blanks. With experiments on multiple languages and datasets, we show that multi-blank RNN-T methods could bring relative speedups of over +90%/+139% to model inference for English Librispeech and German Multilingual Librispeech datasets, respectively. The multi-blank RNN-T method also improves ASR accuracy consistently. We will release our implementation of the method in the NeMo (https://github.com/NVIDIA/NeMo) toolkit. Hainan Xu, Fei Jia, Somshubra Majumdar, Shinji Watanabe 0001, Boris Ginsburg |
ICASSP | 2 |
| 2023 | Efficient Sequence Transduction by Jointly Predicting Tokens and DurationsabstractThis paper introduces a novel Token-and-Duration Transducer (TDT) architecture for sequence-to-sequence tasks. TDT extends conventional RNN-Transducer architectures by jointly predicting both a token and its duration, i.e. the number of input frames covered by the emitted token. This is achieved by using a joint network with two outputs which are independently normalized to generate distributions over tokens and durations. During inference, TDT models can skip input frames guided by the predicted duration output, which makes them significantly faster than conventional Transducers which process the encoder output frame by frame. TDT models achieve both better accuracy and significantly faster inference than conventional Transducers on different sequence transduction tasks. TDT models for Speech Recognition achieve better accuracy and up to 2.82X faster inference than conventional Transducers. TDT models for Speech Translation achieve an absolute gain of over 1 BLEU on the MUST-C test compared with conventional Transducers, and its inference is 2.27X faster. In Speech Intent Classification and Slot Filling tasks, TDT models improve the intent accuracy by up to over 1% (absolute) over conventional Transducers, while running up to 1.28X faster. Our implementation of the TDT model will be open-sourced with the NeMo (https://github.com/NVIDIA/NeMo) toolkit. Hainan Xu, Fei Jia, Somshubra Majumdar, He Huang 0012, Shinji Watanabe 0001, Boris Ginsburg |
ICML | 2 |
| 2023 | A Remote Sensing Image Dehazing Network Based On Dark Channel Attention MechanismabstractClear and haze-free remote sensing image is crucial for subsequent processing and application. We propose a new deep learning network based on the dark channel attention mechanism to enhance remote sensing image dehazing performance by combining prior knowledge with deep learning. The network features a parallel cascade structure of attention flow and dark channel prior (DCP) constraint flow. Additionally, edge loss is introduced to supervise the training network and preserve edge information in the image content. Experiments are conducted on both real and synthetic haze datasets, demonstrating that the proposed method effectively removes non-uniformly distributed haze and produces superior results in both qualitative and quantitative analyses. Zongbao Liang, Fei Jia |
IGARSS | 3 |
| 2023 | A Compact End-to-End Model with Local and Global Context for Spoken Language Identification
Fei Jia, Nithin Rao Koluguri, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 1 |
| 2022 | Joint Attention Mechanism Feature Selection for Single Image Reflection SeparationabstractSeparating the reflective component from a single reflected image has long been an essential but challenging task. To solve the single image reflection separation problem, we combine the reflection model with deep learning, and design a single image reflection separation network based on a nonlinear reflection image mixing model. The network proposed in this paper consists of a joint attention mechanism for learning multidimensional features, and an encoder and a corresponding three-branch decoder. The three feature decoder modules have the same structure but different weight parameters to perform feature selection and decoding on each component in the reflection image, and step through the task of reflection image separation. The separation network structure further considers the performance of different reflection types. Experimental results on synthetic public datasets of three different reflection types and a real public dataset show that the method proposed in this paper can effectively separate the reflection components in multiple types of reflection images. Fei Jia, Yongli Ma, Zongbao Liang |
ICPR | 1 |
| 2022 | NeMo Open Source Speaker Diarization System
Taejin Park, Nithin Rao Koluguri, Fei Jia, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 3 |
| 2022 | Lessons from the AdKDD'21 Privacy-Preserving ML ChallengeabstractDesigning data sharing mechanisms providing performance and strong privacy guarantees is a hot topic for the Online Advertising industry. Namely, a prominent proposal discussed under the Improving Web Advertising Business Group at W3C only allows sharing advertising signals through aggregated, differentially private reports of past displays. To study this proposal extensively, an open Privacy-Preserving Machine Learning Challenge took place at AdKDD’21, a premier workshop on Advertising Science with data provided by advertising company Criteo. In this paper, we describe the challenge tasks, the structure of the available datasets, report the challenge results, and enable its full reproducibility. A key finding is that learning models on large, aggregated data in the presence of a small set of unaggregated data points can be surprisingly efficient and cheap. We also run additional experiments to observe the sensitivity of winning methods to different parameters such as privacy budget or quantity of available privileged side information. We conclude that the industry needs either alternate designs for private data sharing or a breakthrough in learning with aggregated data only to keep ad relevance at a reasonable level. Eustache Diemert, Romain Fabre, Alexandre Gilotte, Fei Jia, Basile Leparmentier, Jérémie Mary, Zhonghua Qu, Ugo Tanielian |
WWW | 4 |
| 2022 | Single image dehazing using generative adversarial networks based on an attention mechanismabstractAbstract Most existing image dehazing methods rely on the solution of the atmospheric scattering model or supervised learning based on paired images. However, owing to incomplete prior knowledge and the lack of paired hazy and haze‐free images of the same scenes as training samples, their performances for single image dehazing are unsatisfactory. Here, the authors present an unpaired image learning method based on the attention mechanism for single image dehazing problems. The method uses the constraint transfer learning ability and circulatory structure of CycleGAN to carry out an unsupervised image dehazing task for unpaired data. Considering the complexity of the haze distribution in actual imaging and human visual characteristics, the improved channel attention and domain attention mechanisms are integrated into the network to process different features and different regions non‐uniformly. The experimental results show that the proposed method achieves good results on both synthetic datasets and real hazy images. Yongli Ma, Fei Jia, Weiqing Yan, Zhaowei Liu 0001, Mengying Ni |
IET Image Process. | 3 |
| 2021 | MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity DetectionabstractWe present MarbleNet, an end-to-end neural network for Voice Activity Detection (VAD). MarbleNet is a deep residual network composed from blocks of 1D time-channel separable convolution, batch-normalization, ReLU and dropout layers. When compared to a state-of-the-art VAD model, MarbleNet is able to achieve similar performance with roughly 1/10-th the parameter cost. We further conduct extensive ablation studies on different training methods and choices of parameters in order to study the robustness of MarbleNet in real-world VAD tasks. Fei Jia, Somshubra Majumdar, Boris Ginsburg |
ICASSP | 1 |
| 2019 | Dynamical Rating Prediction with Topic Words of Reviews: A Hierarchical Analysis Approach
Huibing Zhang, Hao Zhong 0007, Qing Yang 0012, Fei Jia, Fang Pan |
CollaborateCom | 4 |
| 2017 | Aspect-based sentiment analysis using ABPCS model and SVMPperf in Chinese reviewsabstractAspect-based sentiment analysis has always been a difficult task since it consists of several core sub-tasks: feature detection, opinion extraction and polarity classification. Consequently, by now there is little work to summarize all of these works together. In this paper, we propose a brand new holistic system, which can deal with all the problems above simultaneously using aspect-based positive center similarity(ABPCS) model. We experiment our system on clothes and hotel domain, and the result shows considerable improvements over state-of-the-art baselines. Yuxiang Bao, Fei Jia, Xiaoli Bai |
IJCNN | 3 |
| 2015 | Unsupervised Web Topic Detection Using A Ranked Clustering-Like Pattern Across Similarity CascadesabstractDespite the massive growth of social media on the Internet, the process of organizing, understanding, and monitoring user generated content (UGC) has become one of the most pressing problems in today's society. Discovering topics on the web from a huge volume of UGC is one of the promising approaches to achieve this goal. Compared with classical topic detection and tracking in news articles, identifying topics on the web is by no means easy due to the noisy, sparse, and less- constrained data on the Internet. In this paper, we investigate methods from the perspective of similarity diffusion, and propose a clustering-like pattern across similarity cascades (SCs). SCs are a series of subgraphs generated by truncating a similarity graph with a set of thresholds, and then maximal cliques are used to capture topics. Finally, a topic-restricted similarity diffusion process is proposed to efficiently identify real topics from a large number of candidates. Experiments demonstrate that our approach outperforms the state-of-the-art methods on three public data sets. Junbiao Pang, Fei Jia, Chunjie Zhang 0001, Weigang Zhang, Qingming Huang |
IEEE Trans. Multim. | 2 |
| 2014 | Web topic detection using a ranked clustering-like pattern across similarity cascadesabstractIn multi-media and social media communities, web topic detection poses two main difficulties that conventional approaches can barely handle: 1) there are large inter-topic variations among web topics; 2) supervised information is rare to identify the real topics. In this paper, we address these problems from the similarity diffusion perspective among objects on web, and present a clustering-like pattern across similarity cascades (SCs). SCs are a series of subgraphs generated by truncating a weighted graph with a set of thresholds, and then maximal cliques are used to describe the topic candidates. Poisson deconvolution is adopted to efficiently identify the real topics from these topic candidates. Experiments demonstrate that our approach outperforms the state-of-the-arts on two datasets. In addition, we report accuracy v.s. false positives per topic (FPPT) curves for performance evaluation. To our knowledge, this is the first complete evaluation of web topic detection at the topic-wise level, and it establishes a new benchmark for this problem. Fei Jia, Junbiao Pang, Weigang Zhang, Guorong Li, Chunjie Zhang 0001, Qingming Huang, Yugui Liu |
ICME | 1 |