VLDB 2026 Research / reviewers in the wild / expert
Kun Fu 0002
dblp:86/2645-2
· DBLP profile ↗
21ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-2305-1017ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 29% Efficient and distributed learning · 29% Generative modeling · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Bioinformatics and computational biology · 80% Smart cities and intelligent transportation · 20% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 100% |
Topics — the 22 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein function prediction |
1.7 | 2 | 2025 | Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs · KDD (2) 2025 ProtCLIP: Function-Informed Protein Multi-Modal Learning · AAAI 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
1.0 | 2 | 2025 | TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection · EMNLP 2025 CompactETA: A Fast Inference System for Travel Time Prediction · KDD 2020 |
Machine learning › Efficient and distributed learning › attention efficiency
attention sparsity |
0.9 | 1 | 2025 | TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection · EMNLP 2025 |
Machine learning › Generative modeling
generative flow networks |
0.9 | 1 | 2025 | Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer · AAAI 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs · KDD (2) 2025 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
0.9 | 1 | 2025 | TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection · EMNLP 2025 |
Bioinformatics and computational biology › protein design
antibody design |
0.9 | 1 | 2025 | Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer · AAAI 2025 |
Bioinformatics and computational biology
protein design |
0.9 | 1 | 2025 | Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer · AAAI 2025 |
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein representation learning |
0.9 | 1 | 2025 | ProtCLIP: Function-Informed Protein Multi-Modal Learning · AAAI 2025 |
Smart cities and intelligent transportation › traffic estimation
travel time estimation |
0.7 | 2 | 2018 | Learning to Estimate the Travel Time · KDD 2018 Multi-task Representation Learning for Travel Time Estimation · KDD 2018 |
Machine learning › Graph learning › graph neural network › attention-based graph neural network
graph attention network |
0.4 | 1 | 2020 | CompactETA: A Fast Inference System for Travel Time Prediction · KDD 2020 |
Machine learning › Graph learning
spatio-temporal graph learning |
0.4 | 1 | 2020 | CompactETA: A Fast Inference System for Travel Time Prediction · KDD 2020 |
Smart cities and intelligent transportation › traffic prediction
travel time prediction |
0.4 | 1 | 2020 | CompactETA: A Fast Inference System for Travel Time Prediction · KDD 2020 |
Data mining › representation learning › graph representation learning
heterogeneous information network embedding |
0.4 | 1 | 2020 | HetETA: Heterogeneous Information Network Embedding for Estimating Time of Arrival · KDD 2020 |
Machine learning › Representation and self-supervised learning › representation learning › joint representation learning
multi-task representation learning |
0.3 | 1 | 2018 | Multi-task Representation Learning for Travel Time Estimation · KDD 2018 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2018 | Learning to Estimate the Travel Time · KDD 2018 |
Computer vision › Vision and language
image captioning |
0.3 | 1 | 2017 | Aligning Where to See and What to Tell: Image Captioning with Region-Based Attention and Scene-Specific Contexts · IEEE Trans. Pattern Anal. Mach. Intell. 2017 |
Machine learning › Deep learning architectures and training › attention mechanism
visual attention |
0.3 | 1 | 2017 | Aligning Where to See and What to Tell: Image Captioning with Region-Based Attention and Scene-Specific Contexts · IEEE Trans. Pattern Anal. Mach. Intell. 2017 |
Natural language and speech › Language models and text generation › neural language model
protein language model |
0.3 | 1 | 2025 | Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs · KDD (2) 2025 |
Machine learning › Efficient and distributed learning › inference efficiency
low-latency inference |
0.1 | 1 | 2020 | CompactETA: A Fast Inference System for Travel Time Prediction · KDD 2020 |
Storage systems › i/o scheduling
disk scheduling |
0.0 | 1 | 2003 | Comprehensive statistical admission control for streaming media servers · ACM Multimedia 2003 |
Multimedia systems and quality of experience
streaming server |
0.0 | 1 | 2003 | Comprehensive statistical admission control for streaming media servers · ACM Multimedia 2003 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 2.6structure denoising · 1.7protein language model · 1.7products of experts · 1.7potts model · 1.7mixture of experts · 1.7contrastive divergence · 1.7paged dot product kernel · 0.9multimodal pretraining · 0.9dynamic token-level selection · 0.9graph attention network · 0.9temporal convolution · 0.4network embedding · 0.4multilayer perceptron · 0.4graph convolution · 0.4wide linear model · 0.3recurrent neural network · 0.3multi-task learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody DesignerabstractAntibodies defend our health by binding to antigens with high specificity and potentiality, primarily relying on the Complementarity-Determining Region (CDR). Yet, current experimental methods of discovering new antibody CDRs are heavily time-consuming. Computational design could alleviate this burden; especially, protein language models have proven quite beneficial in many recent studies. However, most existing models solely focus on antibody potentiality and struggle to encapsulate the diverse range of plausible CDR candidates, limiting their effectiveness in real-world scenarios as binding is only one factor in the multitude of drug-forming criteria. In this paper, we introduce PG-AbD, a framework uniting Generative Flow Networks (GFlowNets) and pretrained Protein Language Models (PLMs) to successfully generate highly potent, diverse and novel antibody candidates. We innovatively construct a Products of Experts (PoE) composed by the global-distribution-modeling PLM and the local-distribution-modeling Potts Model to serve as the reward function of GFlowNet. The joint training paradigm is introduced, where PoE is trained by contrastive divergence with the negative samples generated by GFlowNet, and then guides GFlowNet to sample diverse antibody candidates. We evaluate PG-AbD on extensive antibody design benchmarks. It significantly outperforms existing methods in diversity (13.5% on RabDab, 31.1% on SabDab) while maintaining optimal potential and novelty. Generated antibodies are also found to form stable, regular 3D structures with their corresponding antigens, demonstrating the great potential of PG-AbD to accelerate real-world antibody discovery. Mingze Yin, Hanjing Zhou, Yiheng Zhu 0002, Jialu Wu, Wei Wu 0045, Kun Fu 0002, Zheng Wang 0027, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001 |
AAAI | 7 |
| 2025 | ProtCLIP: Function-Informed Protein Multi-Modal LearningabstractMulti-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model. Hanjing Zhou, Mingze Yin, Wei Wu 0045, Kun Fu 0002, Jintai Chen, Jian Wu 0001, Zheng Wang 0027 |
AAAI | 5 |
| 2025 | TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache SelectionabstractRapid advances in Large Language Models (LLMs) have spurred demand for processing extended context sequences in contemporary applications.However, this progress faces two challenges: performance degradation due to sequence lengths out-of-distribution, and excessively long inference times caused by the quadratic computational complexity of attention.These issues limit LLMs in long-context scenarios.In this paper, we propose Dynamic Token-Level KV Cache Selection (TokenSelect), a training-free method for efficient and accurate long-context inference.TokenSelect builds upon the observation of non-contiguous attention sparsity, using QK dot products to measure per-head KV Cache criticality at tokenlevel.By per-head soft voting mechanism, To-kenSelect selectively involves a few critical KV cache tokens in attention calculation without sacrificing accuracy.To further accelerate To-kenSelect, we design the Selection Cache based on observations of consecutive Query similarity and implemented the efficient Paged Dot Product Kernel, significantly reducing the selection overhead.A comprehensive evaluation of To-kenSelect demonstrates up to 23.84× speedup in attention computation and up to 2.28× acceleration in end-to-end latency, while providing superior performance compared to state-of-theart long-context inference methods. Wei Wu 0045, Zhuoshi Pan, Kun Fu 0002, Chao Wang 0086, Liyi Chen 0001, Yunchu Bai, Tianfu Wang 0002, Zheng Wang 0027, Hui Xiong 0001 |
EMNLP | 3 |
| 2025 | Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMsabstractProteins, as essential biomolecules, play a central role in biological processes, including metabolic reactions and DNA replication. Accurate prediction of their properties and functions is crucial in biological applications. Recent development of protein language models (pLMs) with supervised fine tuning provides a promising solution to this problem. However, the fine-tuned model is tailored for particular downstream prediction task, and achieving general-purpose protein understanding remains a challenge. In this paper, we introduce Structure-Enhanced Protein Instruction Tuning (SEPIT) framework to bridge this gap. Our approach incorporates a novel structure-aware module into pLMs to enrich their structural knowledge, and subsequently integrates these enhanced pLMs with large language models (LLMs) to advance protein understanding. In this framework, we propose a novel instruction tuning pipeline. First, we warm up the enhanced pLMs using contrastive learning and structure denoising. Then, caption-based instructions are used to establish a basic understanding of proteins. Finally, we refine this understanding by employing a mixture of experts (MoEs) to capture more complex properties and functional information with the same number of activated parameters. Moreover, we construct the largest and most comprehensive protein instruction dataset to date, which allows us to train and evaluate the general-purpose protein understanding model. Extensive experiments on both open-ended generation and closed-set answer tasks demonstrate the superior performance of SEPIT over both closed-source general LLMs and open-source LLMs trained with protein knowledge. Wei Wu 0045, Chao Wang 0086, Liyi Chen 0001, Mingze Yin, Yiheng Zhu 0002, Kun Fu 0002, Jieping Ye, Hui Xiong 0001, Zheng Wang 0027 |
KDD (2) | 6 |
| 2022 | CoDriver ETA: Combine Driver Information in Estimated Time of Arrival by Driving Style Learning Auxiliary TaskabstractEstimated time of arrival (ETA) is one of the most important services in intelligent transportation systems (ITS). Precise ETA ensures proper travel scheduling of passengers as well as guarantees efficient decision-making on ride-hailing platforms, which are used by an explosively growing number of people in the past few years. Recently, machine learning-based methods have been widely adopted to solve this time estimation problem and become state-of-the-art. However, they do not well explore the personalization information, as many drivers are short of personalized data and do not have sufficient trajectory data in real applications. This data sparsity problem prevents existing methods from obtaining higher prediction accuracy. In this article, we propose a novel deep learning method to solve this problem. We introduce an auxiliary task to learn an embedding of the personalized driving information under multi-task learning framework. In this task, we discriminatively learn the embedding of driving preference that preserves the historical statistics of driving speed. For this purpose, we adapt the triplet network from face recognition to learn the embedding by constructing triplets in the feature space. This simultaneously learned embedding can effectively boost the prediction accuracy of the travel time. We evaluate our method on two large-scale real-world datasets from Didi Chuxing platform. The extensive experimental results on billions of historical vehicle travel data demonstrate that the proposed method outperforms state-of-the-art algorithms. Kun Fu 0002, Zheng Wang 0010, Donghua Zhou, Kailun Wu, Jieping Ye, Changshui Zhang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Alleviating Data Sparsity Problems in Estimated Time of Arrival via Auxiliary Metric LearningabstractWith millions of people using ride-hailing platforms for daily travel, estimated time of arrival (ETA) has become a significant problem in intelligent transportation systems and attracted considerable attention recently. Deep learning-based ETA methods have achieved promising results using massive spatial-temporal data. However, we find that the prediction accuracy is not satisfactory in practical applications due to the prevalent data sparsity problems. Instead of focusing on the average prediction performance as many other methods, this study aims to alleviate the data sparsity problems in ETA to enhance user experience. In general, the data sparsity problems arise from two aspects. The first is the road network, where many links are only traversed by few floating cars. The second aspect is drivers, where many drivers’ trajectories are too scarce (e.g., with only 3 trip records). To alleviate the sparsity in road network, we propose a Road Network Metric Learning framework for ETA (RNML-ETA), where an auxiliary metric learning task is used to improve the link-embedding, especially for links with insufficient data. A novel triangle loss is proposed to improve metric learning effectiveness for links. Experiments on massive real-world data show that RNML-ETA outperforms competing methods by promoting the cold links with limited data. Furthermore, we propose a novel unified framework to Alleviate Data Sparsity problems in ETA (ADS-ETA) by extending RNML-ETA with an additional auxiliary task for driver ID embedding. Results with extensive experiments demonstrate that ADS-ETA can effectively alleviate the data sparsity problems caused by road network and driver sparsity. Wenzheng Hu, Donghua Zhou, Baichuan Mo, Kun Fu 0002, Zhengping Che, Zheng Wang 0010, Shenhao Wang, Jinhua Zhao 0001, Jieping Ye, Jian Tang 0008, Changshui Zhang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | FMA-ETA: Estimating Travel Time Entirely Based on FFN with AttentionabstractEstimated time of arrival (ETA) is one of the most important services in intelligent transportation systems (ITS) and becomes a challenging spatial-temporal (ST) data mining task in recent years. Nowadays, deep learning based methods, specifically recurrent neural networks (RNN) based ones are adapted to model the ST patterns from massive data for ETA and become the state-of-the-art. However, RNN is suffering from slow training and inference speed, as its structure is unfriendly to parallel computing. To solve this problem, we propose a novel, brief and effective framework mainly based on feed-forward network (FFN) for ETA, FFN with Multifactor Attention (FMA-ETA). The novel Multi-factor Attention mechanism is proposed to deal with different category features and aggregate the information purposefully. Extensive experimental results on the real-world vehicle travel dataset show FMA-ETA is competitive with state-of-the-art methods in terms of the prediction accuracy with significantly better inference speed. Yulu Wang, Kun Fu 0002, Zheng Wang 0010, Ziang Yan, Changshui Zhang, Jieping Ye |
ICASSP | 3 |
| 2020 | Road Network Metric Learning for Estimated Time of ArrivalabstractRecently, deep learning have achieved promising results in Estimated Time of Arrival (ETA), which is considered as predicting the travel time from the origin to the destination along a given path. One of the key techniques is to use embedding vectors to represent the elements of road network, such as the links (road segments). However, the embedding suffers from the data sparsity problem that many links in the road network are traversed by too few floating cars even in large ride-hailing platforms like Uber and DiDi. Insufficient data makes the embedding vectors in an under-fitting status, which undermines the accuracy of ETA prediction. To address the data sparsity problem, we propose the Road Network Metric Learning framework for ETA (RNML-ETA). It consists of two components: (1) a main regression task to predict the travel time, and (2) an auxiliary metric learning task to improve the quality of link embedding vectors. We further propose the triangle loss, a novel loss function to improve the efficiency of metric learning. We validated the effectiveness of RNML-ETA on large scale realworld datasets, by showing that our method outperforms the state-of-the-art model and the promotion concentrates on the cold links with few data. Kun Fu 0002, Zheng Wang 0010, Changshui Zhang, Jieping Ye |
ICPR | 2 |
| 2020 | Constructing Geographic and Long-term Temporal Graph for Traffic ForecastingabstractTraffic forecasting influences various intelligent transportation system (ITS) services and is of great significance for user experience as well as urban traffic control. It is challenging due to the fact that the road network contains complex and time-varying spatial-temporal dependencies. Recently, deep learning based methods have achieved promising results by adopting graph convolutional network (GCN) to extract the spatial correlations and recurrent neural network (RNN) to capture the temporal dependencies. However, the existing methods often construct the graph only based on road network connectivity, which limits the interaction between roads. In this work, we propose Geographic and Long-term Temporal Graph Convolutional Recurrent Neural Network (GLT-GCRNN), a novel framework for traffic forecasting that learns the rich interactions between roads sharing similar geographic or longterm temporal patterns. Extensive experiments on a real-world traffic state dataset validate the effectiveness of our method by showing that GLT-GCRNN outperforms the state-of-the-art methods in terms of different metrics. Yulu Wang, Kun Fu 0002, Zheng Wang 0010, Changshui Zhang, Jieping Ye |
ICPR | 3 |
| 2020 | CompactETA: A Fast Inference System for Travel Time PredictionabstractComputing estimated time of arrival (ETA) is one of the most important services for online ride-hailing platforms like DiDi and Uber. With billions of service queries per day on such platforms, a fast inference ETA module ensures the efficiency of the overall decision system to guarantee satisfied user experience, as well as saving significant operating cost. In this paper, we develop a novel ETA learning system named as CompactETA, which provides an accurate online travel time inference within 100 microseconds. In the proposed method, we encode high order spatial and temporal dependency into sophisticated representations by applying graph attention network on a spatiotemporal weighted road network graph. We further encode the sequential information of the travel route by positional encoding to avoid the recurrent network structure. The properly learnt representations enable us to apply a very simple multi-layer perceptron model for online real-time inference. Evaluation of both offline experiments and online A/B testing verifies that CompactETA reduces the inference latency by more than 100 times compared to a state-of-the-art system, while maintains competing prediction accuracy. Kun Fu 0002, Fan-Lin Meng, Jieping Ye, Zheng Wang 0010 |
KDD | 1 |
| 2020 | HetETA: Heterogeneous Information Network Embedding for Estimating Time of ArrivalabstractThe estimated time of arrival (ETA) is a critical task in the intelligent transportation system, which involves the spatiotemporal data. Despite a significant amount of prior efforts have been made to design efficient and accurate systems for ETA task, few of them take structural graph data into account, much less the heterogeneous information network. In this paper, we propose HetETA to leverage heterogeneous information graph in ETA task. Specifically, we translate the road map into a multi-relational network and introduce a vehicle-trajectories based network to jointly consider the traffic behavior pattern. Moreover, we employ three components to model temporal information from recent periods, daily periods and weekly periods respectively. Each component comprises temporal convolutions and graph convolutions to learn representations of the spatiotemporal heterogeneous information for ETA task. Experiments on large-scale datasets illustrate the effectiveness of the proposed HetETA beyond the state-of-the-art methods, and show the importance of representation learning of heterogeneous information networks for ETA task. Huiting Hong, Yucheng Lin, Zang Li, Kun Fu 0002, Zheng Wang 0010, Xiaohu Qie, Jieping Ye |
KDD | 5 |
| 2019 | Image Captioning with Partially Rewarded Imitation LearningabstractCurrent state-of-the-art image captioning algorithms have achieved great progress via reinforcement learning or generative adversarial nets, with hand-craft metrics such as CIDEr as the reward for the former and signals from adversarial discriminative networks for the latter. Despite the high scores on metrics or improvement in diversity gained from the application of these methods, they suffer from distinction with human-written sentences and drop of ratings on metrics respectively.In this paper, we propose a novel training objective for image captioning that consists of two parts representing explicit and implicit knowledge respectively. Optimizing the new reward partially with imitation learning, we devise an algorithm in which the caption generator is trained to maximize the combination of CIDEr and predictions from adversarial discriminator. Experiments on MSCOCO dataset demonstrate that the proposed method can integrate the strengths of state-of-the-arts, producing more human-like captions while maintaining comparable performance on traditional metrics. Xintong Yu 0002, Tszhang Guo, Kun Fu 0002, Changshui Zhang, Jianwei Zhang 0001 |
IJCNN | 3 |
| 2018 | Multi-task Representation Learning for Travel Time EstimationabstractOne crucial task in intelligent transportation systems is estimating the duration of a potential trip given the origin location, destination location as well as the departure time. Most existing approaches for travel time estimation assume that the route of the trip is given, which does not hold in real-world applications since the route can be dynamically changed due to traffic conditions, user preferences, etc. As inferring the path from the origin and the destination can be time-consuming and nevertheless error-prone, it is desirable to perform origin-destination travel time estimation, which aims to predict the travel time without online route information. This problem is challenging mainly due to its limited amount of information available and the complicated spatiotemporal dependency. In this paper, we propose a MUlti-task Representation learning model for Arrival Time estimation (MURAT). This model produces meaningful representation that preserves various trip properties in the real-world and at the same time leverages the underlying road network and the spatiotemporal prior knowledge. Further-more, we propose a multi-task learning framework to utilize the path information of historical trips during the training phase which boosts the performance. Experimental results on two large-scale real-world datasets show that the proposed approach achieves clear improvements over state-of-the-art methods Kun Fu 0002, Zheng Wang 0010, Cyrus Shahabi, Jieping Ye, Yan Liu 0002 |
KDD | 2 |
| 2018 | Learning to Estimate the Travel TimeabstractVehicle travel time estimation or estimated time of arrival (ETA) is one of the most important location-based services (LBS). It is becoming increasingly important and has been widely used as a basic service in navigation systems and intelligent transportation systems. This paper presents a novel machine learning solution to predict the vehicle travel time based on floating-car data. First, we formulate ETA as a pure spatial-temporal regression problem based on a large set of effective features. Second, we adapt different existing machine learning models to solve the regression problem. Furthermore, we propose a Wide-Deep-Recurrent (WDR) learning model to accurately predict the travel time along a given route at a given departure time. We then jointly train wide linear models, deep neural networks and recurrent neural networks together to take full advantages of all three models. We evaluate our solution offline with millions of historical vehicle travel data. We also deploy the proposed solution on Didi Chuxing's platform, which services billions of ETA requests and benefits millions of customers per day. Our extensive evaluations show that our proposed deep learning algorithm significantly outperforms the state-of-the-art learning algorithms, as well as the solutions provided by leading industry LBS providers. Zheng Wang 0010, Kun Fu 0002, Jieping Ye |
KDD | 2 |
| 2018 | Image-Text Surgery: Efficient Concept Learning in Image Captioning by Generating PseudopairsabstractImage captioning aims to generate natural language sentences to describe the salient parts of a given image. Although neural networks have recently achieved promising results, a key problem is that they can only describe concepts seen in the training image-sentence pairs. Efficient learning of novel concepts has thus been a topic of recent interest to alleviate the expensive manpower of labeling data. In this paper, we propose a novel method, Image-Text Surgery, to synthesize pseudoimage-sentence pairs. The pseudopairs are generated under the guidance of a knowledge base, with syntax from a seed data set (i.e., MSCOCO) and visual information from an existing large-scale image base (i.e., ImageNet). Via pseudodata, the captioning model learns novel concepts without any corresponding human-labeled pairs. We further introduce adaptive visual replacement, which adaptively filters unnecessary visual features in pseudodata with an attention mechanism. We evaluate our approach on a held-out subset of the MSCOCO data set. The experimental results demonstrate that the proposed approach provides significant performance improvements over state-of-the-art methods in terms of F1 score and sentence quality. An ablation study and the qualitative results further validate the effectiveness of our approach. Kun Fu 0002, Jin Li 0020, Junqi Jin, Changshui Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Aligning Where to See and What to Tell: Image Captioning with Region-Based Attention and Scene-Specific ContextsabstractRecent progress on automatic generation of image captions has shown that it is possible to describe the most salient information conveyed by images with accurate and meaningful sentences. In this paper, we propose an image captioning system that exploits the parallel structures between images and sentences. In our model, the process of generating the next word, given the previously generated ones, is aligned with the visual perception experience where the attention shifts among the visual regions-such transitions impose a thread of ordering in visual perception. This alignment characterizes the flow of latent meaning, which encodes what is semantically shared by both the visual scene and the text description. Our system also makes another novel modeling contribution by introducing scene-specific contexts that capture higher-level semantic information encoded in an image. The contexts adapt language models for word generation to specific scene types. We benchmark our system and contrast to published results on several popular datasets, using both automatic evaluation metrics and human evaluation. We show that either region-based attention or scene-specific contexts improves systems without those components. Furthermore, combining these two modeling ingredients attains the state-of-the-art performance. Kun Fu 0002, Junqi Jin, Runpeng Cui, Fei Sha, Changshui Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Traffic Sign Recognition With Hinge Loss Trained Convolutional Neural NetworksabstractTraffic sign recognition (TSR) is an important and challenging task for intelligent transportation systems. We describe the details of our model's architecture for TSR and suggest a hinge loss stochastic gradient descent (HLSGD) method to train convolutional neural networks (CNNs). Our CNN consists of three stages (70–110–180) with 1 162 284 trainable parameters. The HLSGD is evaluated on the German Traffic Sign Recognition Benchmark, which offers a faster and more stable convergence and a state-of-the-art recognition rate of 99.65%. We write a graphics processing unit package to train several CNNs and establish the final classifier in an ensemble way. Junqi Jin, Kun Fu 0002, Changshui Zhang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2006 | A multi-threshold online smoothing technique for variable rate multimedia streams
Roger Zimmermann, Cyrus Shahabi, Kun Fu 0002, Mehrdad Jahangiri |
Multim. Tools Appl. | 3 |
| 2005 | Randomized Data Allocation in Scalable Streaming Architectures
Kun Fu 0002, Roger Zimmermann |
DASFAA | 1 |
| 2005 | Scalability evaluation of the Yima streaming media architectureabstractAbstract Over the last decade research has been pursued on all aspects of streaming media. While many theoretical results have been reported in the literature, few performance results of large‐scale systems have been published. In this report we specifically explore the scalability aspects of our Yima streaming media architecture in an end‐to‐end test environment. With Yima, it was our goal to design and implement an architecture that would scale in performance from small to large systems. Some of the design features include (1) a multi‐node cluster architecture based on commodity hardware and custom software, (2) media type independence (support ranges from 500 Kb s$^{-1}$ MPEG‐4 to 45 Mb s$^{-1}$ HDTV, at both variable and constant bitrates), (3) fine‐grained online scale up/down capabilities, and (4) a client‐controlled rate smoothing protocol. We briefly discuss the design and implementation of these capabilities of Yima and then thoroughly evaluate its scalability through several sets of experiments. Our results show that Yima scales linearly (within the range of our test parameters) as a function of the cluster size and also as a function of available resources such as network bandwidth and CPU performance. Copyright © 2004 John Wiley & Sons, Ltd. Roger Zimmermann, Cyrus Shahabi, Kun Fu 0002, Shu-Yuen Didi Yao |
Softw. Pract. Exp. | 3 |
| 2003 | Comprehensive statistical admission control for streaming media serversabstractStreaming media servers and digital continuous media recorders require the scheduling of I/O requests to disk drives in real time. There are two accepted paradigms to achieve this: deterministic or statistical. The deterministic approach must assume larger bounds on such disk parameters as the seek time, the rotational latency and the transfer rate, to guarantee the timely service of I/O requests. The statistical approach generally allows higher utilization of resources, in exchange for a residual probability of missed I/O request deadlines. We propose a novel statistical admission control algorithm called TRAC based on a comprehensive three random variable (3RV) model to support both reading and writing of multiple variable bit rate media streams on current generation disk drives. Its major distinctions from previous work include (1) a very realistic disk model which considers multi-zoning of disks, seek and rotational latency profiles, and unequal reading and writing data rate limits, (2) a dynamic bandwidth sharing mechanism between reading and writing, and (3) support for random placement of data blocks. We evaluate the TRAC algorithm through an extensive numerical analysis and real device measurements. The results show that it achieves a much more realistic resource utilization (up to 38\% higher) as compared with the best, previously proposed algorithm based on a single random variable (1RV) model. Most impressive, in all the experiments the difference between the results generated by TRAC and the actual disk device measurements match closely. Roger Zimmermann, Kun Fu 0002 |
ACM Multimedia | 2 |