EDBT 2026 Demo / reviewers in the wild / expert
Jun Rao
dblp:r/JunRao
· DBLP profile ↗
46ranked-venue papers
16as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 25 · 11 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring and enhancing the transfer of distribution in knowledge distillation for autoregressive language models
Jun Rao, Xuebo Liu 0002, Zepeng Lin, Liang Ding 0006, Jing Li 0034, Min Zhang 0005 |
Knowl. Based Syst. | 1 |
| 2026 | Orchestrating Prompt Expertise: Enhancing Knowledge Distillation via Expert-Guided TuningabstractMulti-teacher knowledge distillation transfers knowledge from multiple large teacher models to a small student model and has performed well on many downstream tasks. However, when distilling knowledge from multiple teachers, it always suffers from the severe problems of being time-consuming and storage-extensive for multiple teacher models training and inference. We present MoE-KD, a simple but effective framework that produces supervision for training the student model from one single teacher model, which fixes the above problems and improves effectiveness. In the proposed MoE-KD, multiple trainable prompts are used to extract different views of samples from a single pre-trained language model and only a few parameters (prompts) need to be trained and stored. To guarantee the generated supervision signals with increased robustness and correctness, we introduce an uncertainty-based mechanism and a selector module, which routes the input instance to its corresponding teacher. We have also extended MoE KD to lifelong learning scenarios, proposing a lightweight solution for catastrophic forgetting. We conduct experiments on traditional KD scenarios and lifelong learning scenarios. MoE-KD yields improvements up to 1.1% and 140% in accuracy and efficiency in knowledge distillation and 2.8% improvements on average in lifelong learning, compared with the strong baseline methods. Jun Rao, Shuhan Qi, Xuebo Liu 0002, Lei Wang 0203, Min Zhang 0005, Xuan Wang 0002 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2025 | AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMsabstractDespite the impressive performance of large language models (LLMs) in general domains, they often underperform in specialized domains.Existing approaches typically rely on data synthesis methods and yield promising results by using unlabeled data to capture domain-specific features.However, these methods either incur high computational costs or suffer from performance limitations, while also demonstrating insufficient generalization across different tasks.To address these challenges, we propose AQuilt, a framework for constructing instruction-tuning data for any specialized domains from corresponding unlabeled data, including Answer, Question, Unlabeled data, Inspection, Logic, and Task type.By incorporating logic and inspection, we encourage reasoning processes and self-inspection to enhance model performance.Moreover, customizable task instructions enable high-quality data generation for any task.As a result, we construct a dataset of 703k examples to train a powerful data synthesis model.Experiments show that AQuilt is comparable to DeepSeek-V3 while utilizing just 17% of the production cost.Further analysis demonstrates that our generated data exhibits higher relevance to downstream tasks. Xiaopeng Ke, Hexuan Deng, Xuebo Liu 0002, Jun Rao, Zhenxi Song, Jun Yu 0002, Min Zhang 0005 |
EMNLP | 4 |
| 2025 | A Text-Guided Graph Neural Network for Predicting Long-Range Trajectories of Sea-Surface ObjectsabstractThe Sea-surface environment is influenced by various factors such as wind, waves, tides, and ocean currents. Interactions among various factors, combined with heterogeneous spatio-temporal distribution, lead to highly uncertain dynamics for sea-surface objects. Accordingly, forecasting sea-surface object trajectories constitutes a highly significant challenge. A common approach is to employ temporal prediction networks to model the trajectories of objects in complex motion. However, these methods do not accurately capture the motion state of an object at any given moment. To address this challenge, we propose a Spatial Representation and Motion Pattern Graph Neural Network (SRMPGNN). We represent the motion states of objects as tokenized textual descriptions and perform spatially aligned mapping and interaction with graph neural features. Moreover, long-horizon sea-surface objects are extremely faint, and we propose Graph Super-Resolution Reconstruction (GSR) to address this challenge. We further design spatial representations and Multi-Frequency Enhancement (ME) to highlight the trajectory characteristics of moving objects on the sea-surface. Through end-to-end training, the SRMPGNN achieves state-of-the-art (SOTA) performance on two benchmarks. Yuanlin Zhao, Wei Li 0120, Xiangjun Song, Jun Rao, Jiangang Ding 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | 3AM: An Ambiguity-Aware Multi-Modal Machine Translation DatasetabstractMultimodal machine translation (MMT) is a challenging task that seeks to improve translation quality by incorporating visual information. However, recent studies have indicated that the visual information provided by existing MMT datasets is insufficient, causing models to disregard it and overestimate their capabilities. This issue presents a significant obstacle to the development of MMT research. This paper presents a novel solution to this issue by introducing 3AM, an ambiguity-aware MMT dataset comprising 26,000 parallel sentence pairs in English and Chinese, each with corresponding images. Our dataset is specifically designed to include more ambiguity and a greater variety of both captions and images than other MMT datasets. We utilize a word sense disambiguation model to select ambiguous data from vision-and-language datasets, resulting in a more challenging dataset. We further benchmark several state-of-the-art MMT models on our proposed dataset. Experimental results show that MMT models trained on our dataset exhibit a greater ability to exploit visual information than those trained on other MMT datasets. Our work provides a valuable resource for researchers in the field of multimodal learning and encourages further exploration in this area. The data, code and scripts are freely available at https://github.com/MaxyLee/3AM. Xuebo Liu 0002, Derek F. Wong, Jun Rao, Liang Ding 0006, Lidia S. Chao, Dacheng Tao, Min Zhang 0005 |
LREC/COLING | 4 |
| 2024 | Curriculum Consistency Learning for Conditional Sentence GenerationabstractConsistency learning (CL) has proven to be a valuable technique for improving the robustness of models in conditional sentence generation (CSG) tasks by ensuring stable predictions across various input data forms.However, models augmented with CL often face challenges in optimizing consistency features, which can detract from their efficiency and effectiveness.To address these challenges, we introduce Curriculum Consistency Learning (CCL), a novel strategy that guides models to learn consistency in alignment with their current capacity to differentiate between features.CCL is designed around the inherent aspects of CL-related losses, promoting task independence and simplifying implementation.Implemented across four representative CSG tasks, including instruction tuning (IT) for large language models and machine translation (MT) in three modalities (text, speech, and vision), CCL demonstrates marked improvements.Specifically, it delivers +2.0 average accuracy point improvement compared with vanilla IT and an average increase of +0.7 in COMET scores over traditional CL methods in MT tasks.Our comprehensive analysis further indicates that models utilizing CCL are particularly adept at managing complex instances, showcasing the effectiveness and efficiency of CCL in improving CSG models.Code and scripts are available at https://github.com/xinxinxing/ Curriculum-Consistency-Learning. Liangxin Liu, Xuebo Liu 0002, Lian Lian, Shengjun Cheng, Jun Rao, Tengfei Yu, Hexuan Deng, Min Zhang 0005 |
EMNLP | 5 |
| 2024 | CommonIT: Commonality-Aware Instruction Tuning for Large Language Models via Data PartitionsabstractWith instruction tuning, Large Language Models (LLMs) can enhance their ability to adhere to commands.Diverging from most works focusing on data mixing, our study concentrates on enhancing the model's capabilities from the perspective of data sampling during training.Drawing inspiration from the human learning process, where it is generally easier to master solutions to similar topics through focused practice on a single type of topic, we introduce a novel instruction tuning strategy termed Com-monIT: Commonality-aware Instruction Tuning.Specifically, we cluster instruction datasets into distinct groups with three proposed metrics (TASK, EMBEDDING and LENGTH).We ensure each training mini-batch, or "partition", consists solely of data from a single group, which brings about both data randomness across mini-batches and intra-batch data similarity.Rigorous testing on LLaMa models demonstrates CommonIT's effectiveness in enhancing the instruction-following capabilities of LLMs through IT datasets (FLAN, CoT, and Alpaca) and models (LLaMa2-7B, Qwen2-7B, LLaMa 13B, and BLOOM 7B).CommonIT consistently boosts an average improvement of 2.1% on the general domain (i.e., the average score of Knowledge, Reasoning, Multilinguality and Coding) with the LENGTH metric, and 5.2% on the special domain (i.e., GSM, Openfunctions and Code) with the TASK metric, and 3.8% on the specific tasks (i.e., MMLU) with the EMBEDDING metric.Code is available at https://github.com/raojay7/CommonIT. Jun Rao, Xuebo Liu 0002, Lian Lian, Shengjun Cheng, Yunjie Liao, Min Zhang 0005 |
EMNLP | 1 |
| 2024 | A Novel Series-Concatenation Hybrid Prediction Model of Energy Consumption in Hot Strip Roughing Process With Multi-Step RollingabstractThe steel industry has received serious attention under the background of carbon neutralization and carbon peaking. However, the traditional end-to-end energy consumption (EC) prediction method does not consider the effects of multi-step rolling, multi-time series, and error accumulation. To this end, a series-concatenation model based on a two-stage hybrid network is proposed to achieve EC high-precision prediction in multi-step continuous rolling. Specifically. First, this paper analyzed the mechanism between rolling EC and multiple variables. Second, a two-stage hybrid network with a deep neural network block, convolutional neural network block, long short-term memory block, and SoftBoost block (DCLS-Net) is established for EC prediction of single-step rolling. SoftBoost block is a double-layer structure method proposed in this paper for multi-block precision improvement based on two kinds of boosting strategies. Last, according to the error mechanism and multi-time series characteristics in the multi-step rolling, a series-concatenation EC prediction model is designed to suppress errors and achieve high-precision prediction. The experimental results show that the SoftBoost can effectively improve prediction performance. And the prediction precision of the two-stage DCLS-Net for single-step rolling is improved by 9.43% on average compared with the end-to-end machine learning algorithm. Furthermore, the precision of the series-concatenation model for multi-step rolling is improved by 4.96% compared with the traditional series model, which can satisfy the requirements of high accuracy and positive error in strip rolling production.Note to Practitioners—The inspiration for this paper mainly comes from the multi-objective process planning problem in the hot multi-step continuous roughing process, especially the energy consumption target. This method is also applicable to other variable prediction problems in the continuous multi-step process industry. The common method of energy consumption prediction is end-to-end machine learning. However, the characteristics of multi-step rolling and multi-time series, as well as the influence of intermediate process variables on the results, are not considered, which makes the precision and reliability of prediction unable to satisfy the requirements of industrial production. In this paper, the series-concatenation model structure is adopted to realize the high-precision energy consumption prediction and error suppression of multi-step rolling of strip steel. Such research is helpful for engineers to understand the energy consumption mechanism of strip rolling. It provides guidance for multi-objective roughing process planning. In future research, we will study the influence of more process variables on energy consumption. Yanjiu Zhong, Jun Rao, Shunyu Wu |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Prediction of Energy Consumption in Horizontal Roughing Process of Hot Rolling Strip Based on TDADE AlgorithmabstractThe steel industry is the key industry of energy consumption. The optimization of rolling process parameters is an effective measure to reduce and optimize energy consumption. Accurate energy consumption prediction has an indispensable guiding function in the planning of process parameters. However, the traditional energy consumption prediction and machine learning methods have been unable to fit the needs of high precision and reliability. Taking the horizontal roughing process (HRP) of hot rolling process as the research object, this paper mainly introduces an energy consumption prediction mechanism model (ECPMM) based on a roller adaptive wear strategy by analyzing the strip forming mechanism and rolling process. Then, two directions adaptive differential evolution (TDADE) algorithm is proposed based on a novel adaptive strategy. This algorithm has better convergence speed and global optimization ability than other optimization algorithms in the parameter optimization of ECPMM. We carried out comparative experiments on five machine learning algorithms. The results show that the ECPMM optimized by TDADE is superior to the five machine learning algorithms in prediction accuracy and stability. Finally, we conducted ablation experiments on the energy consumption characteristics of the HRP process, which turns out that using a smaller reduction and roller radius can reduce energy consumption. Summarizing the above experimental results suggests that the ECPMM has high accuracy and robustness, and satisfies the requirements of practical application. Note to Practitioners—It is necessary to carry out the deep analysis and research of energy consumption prediction involved in the process planning of strip hot rolling, the inspiration of this article is stem from this point. The commonly used energy consumption prediction methods include mechanism modeling and machine learning, but both have disadvantages, like low prediction accuracy and black box, which must have to be performed to solve by adapting the production demand of high reliability. In this paper, the combination of the mechanism model and data-driven method is used to realize the accurate prediction of energy consumption and improve the reliability of the model. To the best of our knowledge, this is the first paper about the research on energy consumption prediction of hot rolling process based on mechanism model and data-driven, especially in data size, time granularity, algorithm design, and prediction performance. Such research is helpful to guide engineers to design rolling energy consumption. We will dive a bit deeper into learning the energy consumption and process parameters during the rolling process in future studies. Yanjiu Zhong, Shunyu Wu, Jun Rao, Kangbo Dang |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Parameter-Efficient and Student-Friendly Knowledge DistillationabstractPre-trained models are frequently employed in multimodal learning. However, these models have too many parameters and need too much effort to fine-tune the downstream tasks. Knowledge distillation (KD) is a method to transfer knowledge using the soft label from this pre-trained teacher model to a smaller student, where the parameters of the teacher are fixed (or partially) during training. Recent studies show that this mode may cause difficulties in knowledge transfer due to the mismatched model capacities. To alleviate the mismatch problem, adjustment of temperature parameters, label smoothing and teacher-student joint training methods (online distillation) to smooth the soft label of a teacher network, have been proposed. But those methods rarely explain the effect of smoothed soft labels to enhance the KD performance. The main contributions of our work are the discovery, analysis, and validation of the effect of the smoothed soft label and a less time-consuming and adaptive transfer of the pre-trained teacher's knowledge method, namely PESF-KD by adaptive tuning soft labels of the teacher network. Technically, we first mathematically formulate the mismatch as the sharpness gap between teacher's and student's predictive distributions, where we show such a gap can be narrowed with the appropriate smoothness of the soft label. Then, we introduce an adapter module for the teacher and only update the adapter to obtain soft labels with appropriate smoothness. Experiments on various benchmarks including CV and NLP show that PESF-KD can significantly reduce the training cost while obtaining competitive results compared to advanced online distillation methods. Jun Rao, Xv Meng, Liang Ding 0006, Shuhan Qi, Xuebo Liu 0002, Min Zhang 0005, Dacheng Tao |
IEEE Trans. Multim. | 1 |
| 2024 | Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?abstractThe field of multimedia research has witnessed significant interest in leveraging multimodal pretrained neural network models to perceive and represent the physical world. Among these models, vision-language pretraining (VLP) has emerged as a captivating topic. Currently, the prevalent approach in VLP involves supervising the training process with paired image-text data. However, limited efforts have been dedicated to exploring the extraction of essential linguistic knowledge, such as semantics and syntax, during VLP and understanding its impact on multimodal alignment. In response, our study aims to shed light on the influence of comprehensive linguistic knowledge encompassing semantic expression and syntactic structure on multimodal alignment. To achieve this, we introduce SNARE , a large-scale multimodal alignment probing benchmark designed specifically for the detection of vital linguistic components, including lexical, semantic, and syntax knowledge. SNARE offers four distinct tasks: Semantic Structure, Negation Logic, Attribute Ownership, and Relationship Composition. Leveraging SNARE , we conduct holistic analyses of six advanced VLP models (BLIP, CLIP, Flava, X-VLM, BLIP2, and GPT-4), along with human performance, revealing key characteristics of the VLP model: (i) Insensitivity to complex syntax structures, relying primarily on content words for sentence comprehension. (ii) Limited comprehension of sentence combinations and negations. (iii) Challenges in determining actions or spatial relations within visual information, as well as difficulties in verifying the correctness of ternary relationships. Based on these findings, we propose the following strategies to enhance multimodal alignment in VLP: (1) Utilize a large generative language model as the language backbone in VLP to facilitate the understanding of complex sentences. (2) Establish high-quality datasets that emphasize content words and employ simple syntax, such as short-distance semantic composition, to improve multimodal alignment. (3) Incorporate more fine-grained visual knowledge, such as spatial relationships, into pretraining objectives. 1 Fei Wang 0032, Liang Ding 0006, Jun Rao, Ye Liu 0014, Li Shen 0008, Changxing Ding |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Adaptive Dynamic Programming for Optimal Control of Discrete-Time Nonlinear Systems With Trajectory-Based Initial Control PolicyabstractThe policy gradient adaptive dynamic programming (PGADP) technique has gained recognition as an effective approach for optimizing the performance of nonlinear systems. Nonetheless, existing PGADP algorithms often demand a substantial volume of expensive or potentially risky interaction data with the system. Moreover, the utilization of neural networks in these algorithms can result in suboptimal learning efficiency and unstable training procedures. To address these challenges, a novel algorithm, referred to as OptNet-PGADP, has been introduced. This algorithm integrates an initially tailored control policy based on OptNet to tackle the optimization of control problems in discrete-time nonlinear systems. The OptNet-PGADP algorithm operates through a two-step process. Initially, the input–output trajectory of the system is computed using the nonlinear model predictive control (NMPC) method. Subsequently, an initial admissible control policy is acquired through OptNet. This policy is iteratively enhanced using the PGADP algorithm to attain the optimal controller. The resulting closed-loop control policy can be readily deployed in real-time applications. The implementation of the algorithm employs OptNet for the actor network and integrates an experience replay mechanism to bolster the controller’s learning efficiency. Furthermore, a convergence and optimality analysis of the algorithm is included. Simulation and experimental results conducted on two nonlinear systems conclusively demonstrate that the approach outperforms traditional PGADP and NMPC algorithms. These findings underscore the efficacy of OptNet-PGADP in mitigating the constraints of current methods and achieving superior control performance for nonlinear systems. Jun Rao, Shunyu Wu, Yanjiu Zhong |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Parallel Cross Entropy Policy Gradient Adaptive Dynamic Programming for Optimal Tracking Control of Discrete-Time Nonlinear SystemsabstractPolicy gradient adaptive dynamic programming (PGADP) is a recently acclaimed control technique for the optimal control design of nonlinear systems. Nevertheless, it demands a substantial amount of interaction data with the controlled system, which can prove costly or perilous in certain scenarios. This article introduces a parallel cross entropy optimization method-based PGADP (PCEOM-PGADP) algorithm, with the objective of devising an optimal tracking controller for discrete-time nonlinear systems. The tracking problem is transformed into a regulation problem by constructing a tracking error system. Furthermore, the implementation of the proposed algorithm employs an actor–critic structure, where the actor network represents the control policy and the critic network assesses its performance. Through the iterative interaction, the optimal policy is ultimately derived. The approach also leverages the parallel cross entropy optimization method (PCEOM) to acquire a reasonable initial control policy for PGADP, thereby accelerating the efficiency of the learning process. Convergence analysis of the algorithm is conducted by demonstrating that the generated$Q$function constitutes a monotonically nonincreasing sequence. Finally, the effectiveness of the proposed PCEOM-PGADP algorithm is verified through simulation on a complex automated driving tracking system. Jun Rao, Yanjiu Zhong, Shunyu Wu, Qifang Sun |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Reinforcement Learning Controller Design for Discrete-Time-Constrained Nonlinear Systems With Weight Initialization MethodabstractExtensive research has been dedicated to reinforcement learning (RL) for acquiring proficient optimal controllers through interactions with the environment. However, real-world demands, including enhanced safety performance, introduce considerable challenges to the present design of optimal controllers rooted in RL algorithms. A novel approach is introduced in this article for designing RL-based optimal controllers, employing a control barrier function (CBF) alongside a nonquadratic loss function related to the control signal. The aim is to enable the agent to learn the optimal controller in a secure and efficient manner. To tackle the instability issue in neural network training inherent to traditional RL-based controller design processes, the nonlinear model predictive control (NMPC) technique is employed for initializing the controller network’s weights. A formal demonstration of the method’s optimality is presented. Numerical simulations validate the proposed approach, illustrating its capacity to effectively learn the optimal controller while adhering to the input and state constraints of the system. Yanjiu Zhong, Jun Rao, Shunyu Wu |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2023 | Finetuning Language Models for Multimodal Question AnsweringabstractTo achieve multi-modal intelligence, AI must be able to process and respond to inputs from multimodal sources. However, many current question answering models are limited to specific types of answers, such as yes/no and number, and require additional human assessments. Recently, Visual-Text Question Answering (VQTA) dataset has been proposed to fix this gap. In this paper, we conduct an exhaustive analysis and exploration of this task. Specifically, we implement a T5-based multi-modal generative network that overcomes the limitations of traditional labeling space and provides more freedom in responses. Our approach achieve the best performance in both English and Chinese tracks in the VTQA challenge. Xin Zhang 0097, Wen Xie 0006, Ziqi Dai, Jun Rao, Haokun Wen, Meishan Zhang, Min Zhang 0005 |
ACM Multimedia | 4 |
| 2023 | Data-driven width spread prediction model improvement and parameters optimization in hot strip rolling process
Yanjiu Zhong, Jun Rao, Kangbo Dang |
Appl. Intell. | 4 |
| 2023 | What is the limitation of multimodal LLMs? A deeper look into multimodal LLMs through prompt probing
Shuhan Qi, Zhengying Cao, Jun Rao, Lei Wang 0203, Jing Xiao 0006, Xuan Wang 0002 |
Inf. Process. Manag. | 3 |
| 2023 | Kora: A Cloud-Native Event Streaming Platform for KafkaabstractEvent streaming is an increasingly critical infrastructure service used in many industries and there is growing demand for cloud-native solutions. Confluent Cloud provides a massive scale event streaming platform built on top of Apache Kafka with tens of thousands of clusters running in 70+ regions across AWS, Google Cloud, and Azure. This paper introduces Kora , the cloud-native platform for Apache Kafka at the core of Confluent Cloud. We describe Kora's design that enables it to meet its cloud-native goals, such as reliability, elasticity, and cost efficiency. We discuss Kora's abstractions which allow users to think in terms of their workload requirements and not the underlying infrastructure, and we discuss how Kora is designed to provide consistent, predictable performance across cloud environments with diverse capabilities. Anna Povzner, Prince Mahajan, Jason Gustafson, Jun Rao, Ismael Juma, Feng Min, Shriram Sridharan, Nikhil Bhatia, Gopi K. Attaluri, Adithya Chandra, Stanislav Kozlovski, Rajini Sivaram, Lucas Bradstreet, Bob Barrett, Dhruvil Shah, David Jacot, David Arthur, Manveer Chawla, Ron Dagostino, Colin Mccabe, Manikumar Reddy Obili, Kowshik Prakasam, Jose Garcia Sancio, Alok Nikhil |
Proc. VLDB Endow. | 4 |
| 2023 | Dynamic Contrastive Distillation for Image-Text RetrievalabstractAlthough the vision-and-language pretraining (VLP) equipped cross-modal image-text retrieval (ITR) has achieved remarkable progress in the past two years, it suffers from a major drawback: the ever-increasing size of VLP models restrict its deployment to real-world search scenarios (where the high latency is unacceptable). To alleviate this problem, we present a novel plug-in dynamic contrastive distillation (DCD) framework to compress the large VLP models for the ITR task. Technically, we face the following two challenges: 1) the typical uni-modal metric learning approach is difficult to directly apply to cross-modal task, due to the limited GPU memory to optimize too many negative samples during handling cross-modal fusion features. 2) it is inefficient to static optimize the student network from different hard samples, which have different effects on distillation learning and student network optimization. We try to overcome these challenges from two points. First, to achieve multi-modal contrastive learning, and balance the training costs and effects, we propose to use a teacher network to estimate the difficult samples for students, making the students absorb the powerful knowledge from pre-trained teachers, and master the knowledge from hard samples. Second, to dynamic learn from hard sample pairs, we propose dynamic distillation to dynamically learn samples of different difficulties, from the perspective of better balancing the difficulty of knowledge and students' self-learning ability. We successfully apply our proposed DCD strategy on two state-of-the-art vision-language pretrained models, i.e. ViLT and METER. Extensive experiments on MS-COCO and Flickr 30 K benchmarks show the effectiveness and efficiency of our DCD framework. Encouragingly, we can speed up the inference at least 129 × compared to the existing ITR models. We further provide in-depth analyses and discussions that explain where the performance improvement comes from. We hope our work can shed light on other tasks that require distillation and contrastive learning. Jun Rao, Liang Ding 0006, Shuhan Qi, Yang Liu 0039, Li Shen 0008, Dacheng Tao |
IEEE Trans. Multim. | 1 |
| 2022 | Where Does the Performance Improvement Come From?: - A Reproducibility Concern about Image-Text RetrievalabstractThis article aims to provide the information retrieval community with some reflections on recent advances in retrieval learning by analyzing the reproducibility of image-text retrieval models. Due to the increase of multimodal data over the last decade, image-text retrieval has steadily become a major research direction in the field of information retrieval. Numerous researchers train and evaluate image-text retrieval algorithms using benchmark datasets such as MS-COCO and Flickr30k. Research in the past has mostly focused on performance, with multiple state-of-the-art methodologies being suggested in a variety of ways. According to their assertions, these techniques provide improved modality interactions and hence more precise multimodal representations. In contrast to previous works, we focus on the reproducibility of the approaches and the examination of the elements that lead to improved performance by pretrained and nonpretrained models in retrieving images and text. Jun Rao, Fei Wang 0032, Liang Ding 0006, Shuhan Qi, Yibing Zhan, Weifeng Liu 0001, Dacheng Tao |
SIGIR | 1 |
| 2021 | Student Can Also be a Good Teacher: Extracting Knowledge from Vision-and-Language Model for Cross-Modal RetrievalabstractAstounding results from transformer models with Vision-and Language Pretraining (VLP) on joint vision-and-language downstream tasks have intrigued the multi-modal community. On the one hand, these models are usually so huge that make us more difficult to fine-tune and serve real-time online applications. On the other hand, the compression of the original transformer block will ignore the difference in information between modalities, which leads to the sharp decline of retrieval accuracy. Jun Rao, Tao Qian 0003, Shuhan Qi, Yulin Wu 0001, Qing Liao 0001, Xuan Wang 0002 |
CIKM | 1 |
| 2021 | Consistency and Completeness: Rethinking Distributed Stream Processing in Apache KafkaabstractAn increasingly important system requirement for distributed stream processing applications is to provide strong correctness guarantees under unexpected failures and out-of-order data so that its results can be authoritative (not needing complementary batch results). Although existing systems have put a lot of effort into addressing some specific issues, such as consistency and completeness, how to enable users to make flexible and transparent trade-off decisions among correctness, performance, and cost still remains a practical challenge. Specifically, similar mechanisms are usually applied to tackle both consistency and completeness, which can result in unnecessary performance penalties. We present Apache Kafka's core design for stream processing, which relies on its persistent log architecture as the storage and inter-processor communication layers to achieve correctness guarantees. Kafka Streams, a scalable stream processing client library in Apache Kafka, defines the processing logic as read-process-write cycles in which all processing state updates and result outputs are captured as log appends. Idempotent and transactional write protocols are utilized to guarantee exactly-once semantics. Furthermore, revision-based speculative processing is employed to emit results as soon as possible while handling out-of-order data. We also demonstrate how Kafka Streams behaves in practice with large-scale deployments and performance insights exhibiting its flexible and low-overhead trade-offs. Guozhang Wang, Ayusman Dikshit, Jason Gustafson, Matthias Sax, John Roesler, Sophie Blee-Goldman, Bruno Cadonna, Apurva Mehta, Varun Madan, Jun Rao |
SIGMOD Conference | 12 |
| 2021 | Zone scheduling optimization of pumps in water distribution networks with deep reinforcement learning and knowledge-assisted learning
Jun Rao |
Soft Comput. | 3 |
| 2020 | Internal and Contextual Attention Network for Cold-start Multi-channel Matching in RecommendationabstractReal-world integrated personalized recommendation systems usually deal with millions of heterogeneous items. It is extremely challenging to conduct full corpus retrieval with complicated models due to the tremendous computation costs. Hence, most large-scale recommendation systems consist of two modules: a multi-channel matching module to efficiently retrieve a small subset of candidates, and a ranking module for precise personalized recommendation. However, multi-channel matching usually suffers from cold-start problems when adding new channels or new data sources. To solve this issue, we propose a novel Internal and contextual attention network (ICAN), which highlights channel-specific contextual information and feature field interactions between multiple channels. In experiments, we conduct both offline and online evaluations with case studies on a real-world integrated recommendation system. The significant improvements confirm the effectiveness and robustness of ICAN, especially for cold-start channels. Currently, ICAN has been deployed on WeChat Top Stories used by millions of users. The source code can be obtained from https://github.com/zhijieqiu/ICAN. Ruobing Xie, Zhijie Qiu, Jun Rao, Bo Zhang 0056, Leyu Lin |
IJCAI | 3 |
| 2020 | Group-Aware Long- and Short-Term Graph Representation Learning for Sequential Group RecommendationabstractSequential recommendation and group recommendation are two important branches in the field of recommender system. While considerable efforts have been devoted to these two branches in an independent way, we combine them by proposing the novel sequential group recommendation problem which enables modeling group dynamic representations and is crucial for achieving better group recommendation performance. The major challenge of the problem is how to effectively learn dynamic group representations based on the sequential user-item interactions of group members in the past time frames. To address this, we devise a Group-aware Long- and Short-term Graph Representation Learning approach, namely GLS-GRL, for sequential group recommendation. Specifically, for a target group, we construct a group-aware long-term graph to capture user-item interactions and item-item co-occurrence in the whole history, and a group-aware short-term graph to contain the same information regarding only the current time frame. Based on the graphs, GLS-GRL performs graph representation learning to obtain long-term and short-term user representations, and further adaptively fuse them to gain integrated user representations. Finally, group representations are obtained by a constrained user-interacted attention mechanism which encodes the correlations between group members. Comprehensive experiments demonstrate that GLS-GRL achieves better performance than several strong alternatives coming from sequential recommendation and group recommendation methods, validating the effectiveness of the core components in GLS-GRL. Wen Wang 0016, Wei Zhang 0056, Jun Rao, Zhijie Qiu, Bo Zhang 0056, Leyu Lin, Hongyuan Zha |
SIGIR | 3 |
| 2019 | Road Segmentation Based on Hybrid Convolutional Network for High-Resolution Visible Remote Sensing ImageabstractRoad segmentation plays an important role in many applications, such as intelligent transportation system and urban planning. Various road segmentation methods have been proposed for visible remote sensing images, especially the popular convolutional neural network-based methods. However, high-accuracy road segmentation from high-resolution visible remote sensing images is still a challenging problem due to complex background and multiscale roads in these images. To handle this problem, a hybrid convolutional network (HCN), fusing multiple subnetworks, is proposed in this letter. The HCN contains a fully convolutional network, a modified U-Net, and a VGG subnetwork; these subnetworks obtain a coarse-grained, a medium-grained, and a fine-grained road segmentation map. Moreover, the HCN uses a shallow convolutional subnetwork to fuse these multigrained segmentation maps for final road segmentation. Benefitting from multigrained segmentation, our HCN shows impressing results in processing both multiscale roads and complex background. Four testing indicators, including pixel accuracy, mean accuracy, mean region intersection over union (IU), and frequency weighted IU, are computed to evaluate the proposed HCN on two testing data sets. Compared with five state-of-the-art road segmentation methods, our HCN has higher segmentation accuracy than them. Ye Li 0013, Jun Rao, Lele Xu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Optimal caching placement for D2D assisted wireless caching networksabstractIn this paper, we devise the optimal caching placement to maximize the offloading probability for a two-tier wireless caching system, where the helpers and a part of users have caching ability. The offloading comes from the local caching, D2D sharing and the helper transmission. In particular, to maximize the offloading probability we reformulate the caching placement problem for users and helpers into a difference of convex (DC) problem which can be effectively solved by DC programming. Moreover, we analyze the two extreme cases where there is only help-tier caching network and only user-tier. Specifically, the placement problem for the helper-tier caching network is reduced to a convex problem, and can be effectively solved by the classical water-filling method. We notice that users and helpers prefer to cache popular contents under low node density and prefer to cache different contents evenly under high node density. Simulation results indicate a great performance gain of the proposed caching placement over existing approaches. Jun Rao, Zhiyong Chen 0002, Bin Xia 0001 |
ICC | 1 |
| 2015 | Liquid: Unifying Nearline and Offline Big Data Integration
Raul Castro Fernandez, Peter R. Pietzuch, Jay Kreps, Neha Narkhede, Jun Rao, Joel Koshy, Dong Lin, Chris Riccomini, Guozhang Wang |
CIDR | 5 |
| 2015 | Building a Replicated Logging System with Apache KafkaabstractApache Kafka is a scalable publish-subscribe messaging system with its core architecture as a distributed commit log. It was originally built at LinkedIn as its centralized event pipelining platform for online data integration tasks. Over the past years developing and operating Kafka, we extend its log-structured architecture as a replicated logging backbone for much wider application scopes in the distributed environment. In this abstract, we will talk about our design and engineering experience to replicate Kafka logs for various distributed data-driven systems at LinkedIn, including source-of-truth data storage and stream processing. Guozhang Wang, Joel Koshy, Sriram Subramanian, Kartik Paramasivam, Mammad Zadeh, Neha Narkhede, Jun Rao, Jay Kreps, Joe Stein |
Proc. VLDB Endow. | 7 |
| 2012 | Data Infrastructure at LinkedInabstractLinked In is among the largest social networking sites in the world. As the company has grown, our core data sets and request processing requirements have grown as well. In this paper, we describe a few selected data infrastructure projects at Linked In that have helped us accommodate this increasing scale. Most of those projects build on existing open source projects and are themselves available as open source. The projects covered in this paper include: (1) Voldemort: a scalable and fault tolerant key-value store, (2) Data bus: a framework for delivering database changes to downstream applications, (3) Espresso: a distributed data store that supports flexible schemas and secondary indexing, (4) Kafka: a scalable and efficient messaging system for collecting various user activity events and log data. Aditya Auradkar, Chavdar Botev, Shirshanka Das, Dave De Maagd, Alex Feinberg, Phanindra Ganti, Bhaskar Ghosh, Kishore Gopalakrishna, Brendan Harris, Joel Koshy, Kevin Krawez, Jay Kreps, Shi Lu, Sunil Nagaraj, Neha Narkhede, Sasha Pachev, Igor Perisic, Lin Qiao, Tom Quiggle, Jun Rao, Bob Schulman, Abraham Sebastian, Oliver Seeliger, Adam Silberstein, Boris Shkolnik, Chinmay Soman, Roshan Sumbaly, Kapil Surlaker, Sajid Topiwala, Cuong Tran 0003, Balaji Varadarajan, Jemiah Westerman, Zach White, David Zhang 0001 |
ICDE | 21 |
| 2011 | Using Paxos to Build a Scalable, Consistent, and Highly Available DatastoreabstractSpinnaker is an experimental datastore that is designed to run on a large cluster of commodity servers in a single datacenter. It features key-based range partitioning, 3-way replication, and a transactional get-put API with the option to choose either strong or timeline consistency on reads. This paper describes Spinnaker's Paxos-based replication protocol. The use of Paxos ensures that a data partition in Spinnaker will be available for reads and writes as long a majority of its replicas are alive. Unlike traditional master-slave replication, this is true regardless of the failure sequence that occurs. We show that Paxos replication can be competitive with alternatives that provide weaker consistency guarantees. Compared to an eventually consistent datastore, we show that Spinnaker can be as fast or even faster on reads and only 5% to 10% slower on writes. Jun Rao, Eugene J. Shekita, Sandeep Tata |
Proc. VLDB Endow. | 1 |
| 2010 | A comparison of join algorithms for log processing in MaPreduceabstractThe MapReduce framework is increasingly being used to analyze large volumes of data. One important type of data analysis done with MapReduce is log processing, in which a click-stream or an event log is filtered, aggregated, or mined for patterns. As part of this analysis, the log often needs to be joined with reference data such as information about users. Although there have been many studies examining join algorithms in parallel and distributed DBMSs, the MapReduce framework is cumbersome for joins. MapReduce programmers often use simple but inefficient algorithms to perform joins. In this paper, we describe crucial implementation details of a number of well-known join strategies in MapReduce, and present a comprehensive experimental comparison of these join techniques on a 100-node Hadoop cluster. Our results provide insights that are unique to the MapReduce platform and offer guidance on when to use a particular join algorithm on this platform. Spyros Blanas, Jignesh M. Patel, Vuk Ercegovac, Jun Rao, Eugene J. Shekita, Yuanyuan Tian 0001 |
SIGMOD Conference | 4 |
| 2008 | Dynamic faceted search for discovery-driven analysisabstractWe propose a dynamic faceted search system for discovery-driven analysis on data with both textual content and structured attributes. From a keyword query, we want to dynamically select a small set of "interesting" attributes and present aggregates on them to a user. Similar to work in OLAP exploration, we define "interestingness" as how surprising an aggregated value is, based on a given expectation. We make two new contributions by proposing a novel "navigational" expectation that's particularly useful in the context of faceted search, and a novel interestingness measure through judicious application of p-values. Through a user survey, we find the new expectation and interestingness metric quite effective. We develop an efficient dynamic faceted search system by improving a popular open source engine, Solr. Our system exploits compressed bitmaps for caching the posting lists in an inverted index, and a novel directory structure called a bitset tree for fast bitset intersection. We conduct a comprehensive experimental study on large real data sets and show that our engine performs 2 to 3 times faster than Solr. Debabrata Dash, Jun Rao, Nimrod Megiddo, Anastasia Ailamaki, Guy M. Lohman |
CIKM | 2 |
| 2008 | DBPubs: multidimensional exploration of database publicationsabstractDBPubs is a system for effectively analyzing and exploring the content of database publications by combining keyword search with OLAP-style aggregations, navigation, and reporting. DBPubs starts with keyword search over the content of publications. The publications' metadata such as title, authors, venues, year, and so on, provide traditional OLAP static dimensions, which are combined with dynamic dimensions discovered from the content of the publications in the search result, such as frequent phrases, relevant phrases, and topics. We compute publication ranks based on the link structure between documents, i.e., citations, and aggregate them to find seminal papers, discover trends, and rank authors. We deploy an OLAP tool for multidimensional content exploration through traditional OLAP rollup-drilldown operations on the static and dynamic dimensions, solutions for multi-cube analysis, dynamic navigation of the content, and highlighting of interesting dices of the multidimensional content dataspace. Akanksha Baid, Andrey Balmin, Heasoo Hwang, Erik Nijkamp, Jun Rao, Berthold Reinwald, Alkis Simitsis, Yannis Sismanis, Frank van Ham |
Proc. VLDB Endow. | 5 |
| 2007 | Impliance: A Next Generation Information Management Appliance
Bishwaranjan Bhattacharjee, Joseph S. Glider, Richard A. Golding, Guy M. Lohman, Volker Markl, Hamid Pirahesh, Jun Rao, Robert M. Rees, Garret Swart |
CIDR | 7 |
| 2006 | Compiled Query Execution Engine using JVMabstractA conventional query execution engine in a database system essentially uses a SQL virtual machine (SVM) to interpret a dataflow tree in which each node is associated with a relational operator. During query evaluation, a single tuple at a time is processed and passed among the operators. Such a model is popular because of its efficiency for pipelined processing. However, since each operator is implemented statically, it has to be very generic in order to deal with all possible queries. Such generality tends to introduce significant runtime inefficiency, especially in the context of memory-resident systems, because the granularity of data commercial system, using SVM. processing (a tuple) is too small compared with the associated overhead. Another disadvantage in such an engine is that each operator code is compiled statically, so query-specific optimization cannot be applied. To improve runtime efficiency, we propose a compiled execution engine, which, for a given query, generates new query-specific code on the fly, and then dynamically compiles and executes the code. The Java platform makes our approach particularly interesting for several reasons: (1) modern Java Virtual Machines (JVM) have Just- In-Time (JIT) compilers that optimize code at runtime based on the execution pattern, a key feature that SVMs lack; (2) because of Java’s continued popularity, JVMs keep improving at a faster pace than SVMs, allowing us to exploit new advances in the Java runtime in the future; (3) Java is a dynamic language, which makes it convenient to load a piece of new code on the fly. In this paper, we develop both an interpreted and a compiled query execution engine in a relational, Java-based, in-memory database prototype, and perform an experimental study. Our experimental results on the TPC-H data set show that, despite both engines benefiting from JIT, the compiled engine runs on average about twice as fast as the interpreted one, and significantly faster than an in-memory Jun Rao, Hamid Pirahesh, C. Mohan 0001, Guy M. Lohman |
ICDE | 1 |
| 2006 | A Deferred Cleansing Method for RFID Data Analytics
Jun Rao, Sangeeta Doraiswamy, Hetal Thakkar, Latha S. Colby |
VLDB | 1 |
| 2004 | Canonical Abstraction for Outerjoin OptimizationabstractOuterjoins are an important class of joins and are widely used in various kinds of applications. It is challenging to optimize queries that contain outerjoins because outerjoins do not always commute with inner joins. Previous work has studied this problem and provided techniques that allow certain reordering of the join sequences. However, the optimization of outerjoin queries is still not as powerful as that of inner joins.An inner join query can always be canonically represented as a sequence of Cartesian products of all relations, followed by a sequence of selection operations, each applying a conjunct in the join predicates. This canonical abstraction is very powerful because it enables the optimizer to use any join sequence for plan generation. Unfortunately, such a canonical abstraction for outerjoin queries has not been developed. As a result, existing techniques always exclude certain join sequences from planning, which can lead to a severe performance penalty.Given a query consisting of a sequence of inner and outer joins, we, for the first time, present a canonical abstraction based on three operations: outer Cartesian products, nullification, and best match. Like the inner join abstraction, our outerjoin abstraction permits all join sequences, and preserves the property of both commutativity and transitivity among predicates. This allows us to generate plans that are very desirable for performance reasons but that couldn't be done before. We present an algorithm that produces such a canonical abstraction, and a method that extends an inner-join optimizer to generate plans in an expanded search space. We also describe an efficient implementation of the best match operation using the OLAP functionalities in SQL:1999. Our experimental results show that our technique can significantly improve the performance of outerjoin queries. Jun Rao, Hamid Pirahesh, Calisto Zuzarte |
SIGMOD Conference | 1 |
| 2004 | DB2 Design Advisor: Integrated Automatic Physical Database Design
Daniel C. Zilio, Jun Rao, Sam Lightstone, Guy M. Lohman, Adam J. Storm, Christian Garcia-Arellano, Scott Fadden |
VLDB | 2 |
| 2003 | Estimating Compilation Time of a Query OptimizerabstractA query optimizer compares alternative plans in its search space to find the best plan for a given query. Depending on the search space and the enumeration algorithm, optimizers vary in their compilation time and the quality of the execution plan they can generate. This paper describes a compilation time estimator that provides a quantified estimate of the optimizer compilation time for a given query. Such an estimator is useful for automatically choosing the right level of optimization in commercial database systems. In addition, compilation time estimates can be quite helpful for mid-query reoptimization, for monitoring the progress of workload analysis tools where a large number queries need to be compiled (but not executed), and for judicious design and tuning of an optimizer.Previous attempts to estimate optimizer compilation complexity used the number of possible binary joins as the metric and overlooked the fact that each join often translates into a different number of join plans because of the presence of "physical" properties. We use the number of plans (instead of joins) to estimate query compilation time, and employ two novel ideas: (1) reusing an optimizer's join enumerator to obtain actual number of joins, but bypassing plan generation to save estimation overhead; (2) maintaining a small number of "interesting" properties to facilitate plan counting. We prototyped our approach in a commercial database system and our experimental results show that we can achieve good compilation time estimates (less than 30% error, on average) for complex real queries, using a small fraction (within 3%) of the actual compilation time. Ihab F. Ilyas, Jun Rao, Guy M. Lohman, Dengfeng Gao, Eileen Tien Lin |
SIGMOD Conference | 2 |
| 2002 | Automating physical database design in a parallel databaseabstractPhysical database design is important for query performance in a shared-nothing parallel database system, in which data is horizontally partitioned among multiple independent nodes. We seek to automate the process of data partitioning. Given a workload of SQL statements, we seek to determine automatically how to partition the base data across multiple nodes to achieve overall optimal (or close to optimal) performance for that workload. Previous attempts use heuristic rules to make those decisions. These approaches fail to consider all of the interdependent aspects of query performance typically modeled by today's sophisticated query optimizers.We present a comprehensive solution to the problem that has been tightly integrated with the optimizer of a commercial shared-nothing parallel database system. Our approach uses the query optimizer itself both to recommend candidate partitions for each table that will benefit each query in the workload, and to evaluate various combinations of these candidates. We compare a rank-based enumeration method with a random-based one. Our experimental results show that the former is more effective. Jun Rao, Nimrod Megiddo, Guy M. Lohman |
SIGMOD Conference | 1 |
| 2001 | Using EELs, a Practical Approach to Outerjoin and Antijoin ReorderingabstractOuterjoins and antijoins are two important classes of joins in database systems. Reordering outerjoins and antijoins with innerjoins is challenging because not all the join orders preserve the semantics of the original query. Previous work did not consider antijoins and was restricted to a limited class of queries. We consider using a conventional bottom-up optimizer to reorder different types of joins. We propose extending each join predicate's eligibility list, which contains all the tables referenced in the predicate. An extended eligibility list (EEL) includes all the tables needed by a predicate to preserve the semantics of the original query. We describe an algorithm that can set up the EELs properly in a bottom-up traversal of the original operator tree. A conventional join optimizer is then modified to check the EELs when generating sub-plans. Our approach handles antijoin and can resolve many practical issues. It is now being implemented in an upcoming release of IBM's Universal Database Server for Unix, Windows and OS/2. Jun Rao, Bruce G. Lindsay 0001, Guy M. Lohman, Hamid Pirahesh, David E. Simmen |
ICDE | 1 |
| 2001 | Adapting materialized views after redefinitions: techniques and a performance study
Ashish Gupta 0001, Inderpal Singh Mumick, Jun Rao, Kenneth A. Ross |
Inf. Syst. | 3 |
| 2000 | Making B+-Trees Cache Conscious in Main MemoryabstractPrevious research has shown that cache behavior is important for main memory index structures. Cache conscious index structures such as Cache Sensitive Search Trees (CSS-Trees) perform lookups much faster than binary search and T-Trees. However, CSS-Trees are designed for decision support workloads with relatively static data. Although B+-Trees are more cache conscious than binary search and T-Trees, their utilization of a cache line is low since half of the space is used to store child pointers. Nevertheless, for applications that require incremental updates, traditional B+-Trees perform well. Jun Rao, Kenneth A. Ross |
SIGMOD Conference | 1 |
| 1999 | Cache Conscious Indexing for Decision-Support in Main Memory
Jun Rao, Kenneth A. Ross |
VLDB | 1 |
| 1998 | Reusing Invariants: A New Strategy for Correlated QueriesabstractCorrelated queries are very common and important in decision support systems. Traditional nested iteration evaluation methods for such queries can be very time consuming. When they apply, query rewriting techniques have been shown to be much more efficient. But query rewriting is not always possible. When query rewriting does not apply, can we do something better than the traditional nested iteration methods? In this paper, we propose a new invariant technique to evaluate correlated queries efficiently. The basic idea is to recognize the part of the subquery that is not related to the outer references and cache the result of that part after its first execution. Later, we can reuse the result and combine it with the result of the rest of the subquery that is changing for each iteration. Our technique applies to arbitrary correlated subqueries.This paper introduces algorithms to recognize the invariant part of a data flow tree, and to restructure the evaluation plan to reuse the stored intermediate result. We also propose an efficient method to teach an existing join optimizer to understand the invariant feature and thus allow it to be able to generate better join plans in the new context. Some other related optimization techniques are also discussed. The proposed techniques were implemented within three months on an existing real commercial database system.We also experimentally evaluate our proposed technique. Our evaluation indicates that, when query rewriting is not possible, the invariant technique is significantly better than the traditional nested iteration method. Even when query rewriting applies, the invariant technique is sometimes better than the query rewriting technique. Our conclusion is that the invariant technique should be considered as one of the alternatives in evaluating correlated queries since it fills the gap left by rewriting techniques. Jun Rao, Kenneth A. Ross |
SIGMOD Conference | 1 |