VLDB 2026 Research / reviewers in the wild / expert
Yujia Xie
dblp:201/8729
· DBLP profile ↗
24ranked-venue papers
8as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-Based Large Language ModelsabstractVideo-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user inquiries. In pursuit of the ultimate goal of achieving artificial general intelligence, a truly intelligent Video-LLM model should not only see and understand the surroundings, but also possess human-level commonsense, and make well-informed decisions for users. To guide the development of such a model, the establishment of a robust and comprehensive evaluation system becomes crucial. To this end, this paper proposes Video-Bench, a new comprehensive benchmark along with a toolkit specifically designed for evaluating Video-LLMs. The benchmark comprises 10 meticulously crafted tasks, evaluating the capabilities of Video-LLMs across three distinct levels: video-exclusive understanding, prior knowledge-based question-answering, and comprehension and decision-making. In addition, we introduce an automatic toolkit tailored to process model outputs for various tasks, facilitating the calculation of metrics and conveniently generating final scores. We evaluate 9 representative Video-LLMs using Video-Bench. The findings reveal that current Video-LLMs still fall considerably short of achieving human-like comprehension and analysis of real-world video, and offer valuable insights for future research directions. The benchmark and toolkit are available at https://github.com/PKU-YuanGroup/Video-Bench. Munan Ning, Yujia Xie, Bin Lin 0014, Jiaxi Cui, Lu Yuan 0001, Dongdong Chen 0001, Li Yuan 0007 |
Comput. Vis. Media | 3 |
| 2026 | AMC-GPT: Integrating physics-informed interference emulation into Generative Pre-trained Transformers for AMC
Shuyuan Yang 0001, Zhixi Feng, Yifan Gai, Yujia Xie |
Knowl. Based Syst. | 5 |
| 2025 | A Traceable and Anonymous Mutual Authentication Scheme for Smart Healthcare on Elliptic CurvesabstractABSTRACT The rapid development of big data technologies has exacerbated the challenge of maintaining patient privacy in smart healthcare environments. Although previous mutual patient–physician authentication systems achieve basic anonymization, patients' communication addresses are still exposed, and attackers can analyze transaction records to establish correlations between users' addresses and even obtain their real identities. To address this problem, we propose a user anonymization scheme based on the elliptic curve discrete logarithmic problem assumption, which aims to prevent malicious interception and theft of patients' personal data by obfuscating the identity of registered users. By combining identity‐based encryption with advanced anonymization techniques and reconstructing signatures of knowledge, traceability is achieved while ensuring that only the intended recipient with the corresponding private key can decrypt the data. The validation shows that our system guarantees unlinkability and anonymity while resisting hijacking attacks and man‐in‐the‐middle attacks, and it is simulated using JPBC 2.0.0 (Jdk version 14.0.1), which shows that the communication overhead needs 808 bytes and that the computation overhead for system initialization, signature, and validation are 102, 167, and 70 ms, respectively. Yujia Xie, Wenjing Lv |
Concurr. Comput. Pract. Exp. | 1 |
| 2025 | Tractor Semi-Trailer Off-Tracking and Stability Approximate Bi-Level Policy OptimizationabstractTrajectory tracking control of tractor semi-trailer vehicles poses significant challenges due to inherent off-tracking behavior and roll instability risks. While existing approaches have demonstrated effectiveness, they often rely on computationally intensive numerical solvers and require time-consuming manual tuning of cost function weights. This paper presents an approximate bi-level policy optimization (ABPO) framework that simultaneously optimizes the cost function and synthesizes an explicit control policy to minimize off-tracking while reducing computational complexity. The proposed framework employs a hierarchical structure: the upper level updates cost weights based on the trailer’s stability trajectory, while the lower level derives an approximate optimal policy by solving the tractor’s control problem. By leveraging Pontryagin’s Maximum Principle (PMP), we have developed a novel method to analytically compute cost weight gradients through differentiation of the PMP conditions. This enables the formulation of a related optimal control problem (OCP) whose solutions directly yield gradients for cost parameter updates. The ABPO framework achieves automatic weight coefficient adjustment, enhances trajectory tracking accuracy for both tractor and trailer units, and significantly reduces computational burden. Simulation and experimental validation across 4 classical scenarios demonstrates that the learned policy reduces rearward amplification by 17.82%, lateral tracking errors by 84.15%, and rollover by 64.19%, respectively. Notably, the control policy computation requires less than 10 ms, making it suitable for real-time applications. The source code for the algorithms described in this paper is publicly available at https://github.com/TroyResearch/ABPO.git. Fawang Zhang, Jingliang Duan, Hui Liu 0001, Xingyu Cao, Shida Nie, Congshuai Guo, Yujia Xie, Jun Ma 0008, Shangli Wang |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | Personalized Off-Road Path Planning Based on Internal and External Characteristics for Obstacle AvoidanceabstractOff-road environments with varied terrain and obstacle types present substantial challenges to the safe maneuvering of unmanned ground vehicles (UGVs). This study addresses the need for personalized path planning by introducing a multi-source off-road potential field (MOPF) method that quantifies risk and impediments in off-road settings based on internal and external characteristics. Specifically, Vehicle capability boundaries are defined by longitudinal dynamics analysis of the ego-vehicle to prevent instability due to insufficient driving force and limited adhesion conditions. A novel Non-Uniform Safety Margin Expression (NSME) is proposed to adjust the MOPF, allowing it to consider the vehicle’s state to enhance travel efficiency and minimize detours. The MOPF can be adapted according to the characteristics of the ego vehicle, drivers, and cargo. To incorporate driving styles, the Driving Style Probabilistic Roadmap (DSPRM) algorithm is developed, leading to smoother and more personalized paths. Comparative tests demonstrate that our method enables personalized path planning, achieving an average reduction of 10.29% in path length and 30.83% in path slope compared to traditional planning methods, while maintaining a safe distance from obstacles. Shida Nie, Yujia Xie, Congshuai Guo, Hui Liu 0001, Fawang Zhang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsabstractDespite their impressive capabilities, large language models (LLMs) are prone to hallucinations, i.e., generating content that deviates from facts seen during pretraining. We propose a simple decoding strategy for reducing hallucinations with pretrained LLMs that does not require conditioning on retrieved external knowledge nor additional fine-tuning. Our approach obtains the next-token distribution by contrasting the differences in logits obtained from projecting the later layers versus earlier layers to the vocabulary space, exploiting the fact that factual knowledge in an LLMs has generally been shown to be localized to particular transformer layers. We find that this **D**ecoding by C**o**ntrasting **La**yers (DoLa) approach is able to better surface factual knowledge and reduce the generation of incorrect facts. DoLa consistently improves the truthfulness across multiple choices tasks and open-ended generation tasks, for example improving the performance of LLaMA family models on TruthfulQA by 12-17% absolute points, demonstrating its potential in making LLMs reliably generate truthful facts. Yung-Sung Chuang, Yujia Xie, Hongyin Luo, James R. Glass |
ICLR | 2 |
| 2024 | Identifying Socially Optimal Equilibria Using Combinatorial Properties of Nash Equilibria in Bimatrix GamesabstractNash equilibrium is arguably the most fundamental concept in game theory, which is used to analyze and predict the behavior of the players. In many games, there exist multiple equilibria, with different expected payoffs for the players, which in turn raises the question of equilibrium selection. In this paper, we study the [Formula: see text]-hard problem of identifying a socially optimal Nash equilibrium in two-player normal-form games (called bimatrix games), which may be represented by a mixed integer linear program (MILP). We characterize the properties of the equilibria and develop several classes of valid inequalities accordingly. We use these theoretical results to provide a decomposition-based reformulation of the MILP, which we solve by a branch-and-cut algorithm. Our extensive computational experiments demonstrate superiority of our approach over solving the MILP formulation through feeding it into a commercial solver or through the “traditional” Benders’ decomposition. Of note, our proposed approach can find provably optimal solutions for many instances. History: Accepted by Andrea Lodi, Area Editor for Design & Analysis of Algorithms. Funding: This work was supported by the National Institute of Dental and Craniofacial Research [Grant R01DE028283]. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. The funding agreements ensured the authors’ independence in designing the study, interpreting the data, writing, and publishing the report. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.0072 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2022.0072 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Amin Dehghanian, Yujia Xie, Nicoleta Serban |
INFORMS J. Comput. | 2 |
| 2023 | i-Code: An Integrative and Composable Multimodal Learning FrameworkabstractHuman intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised pretraining framework where users may flexibly combine the modalities of vision, speech, and language into unified and general-purpose vector representations. In this framework, data from each modality are first given to pretrained single-modality encoders. The encoder outputs are then integrated with a multimodal fusion network, which uses novel merge- and co-attention mechanisms to effectively combine information from the different modalities. The entire system is pretrained end-to-end with new objectives including masked modality unit modeling and cross-modality contrastive learning. Unlike previous research using only video for pretraining, the i-Code framework can dynamically process single, dual, and triple-modality data during training and inference, flexibly projecting different combinations of modalities into a single representation space. Experimental results demonstrate how i-Code can outperform state-of-the-art techniques on five multimodal understanding tasks and single-modality benchmarks, improving by as much as 11% and demonstrating the power of integrative multimodal pretraining. Ziyi Yang 0011, Yuwei Fang, Chenguang Zhu 0001, Reid Pryzant, Dongdong Chen 0001, Yu Shi 0001, Yichong Xu, Yao Qian, Mei Gao, Liyang Lu, Yujia Xie, Robert Gmyr, Noel Codella, Naoyuki Kanda, Bin Xiao 0004, Lu Yuan 0001, Takuya Yoshioka, Michael Zeng 0001, Xuedong Huang 0001 |
AAAI | 12 |
| 2023 | Look Before You Match: Instance Understanding Matters in Video Object SegmentationabstractExploring dense matching between the current frame and past frames for long-range context modeling, memory-based methods have demonstrated impressive results in video object segmentation (VOS) recently. Nevertheless, due to the lack of instance understanding ability, the above approaches are oftentimes brittle to large appearance variations or viewpoint changes resulted from the movement of objects and cameras. In this paper, we argue that instance understanding matters in VOS, and integrating it with memory-based matching can enjoy the synergy, which is intuitively sensible from the definition of VOS task, i.e., identifying and segmenting object instances within the video. Towards this goal, we present a two-branch network for VOS, where the query-based instance segmentation (IS) branch delves into the instance details of the current frame and the VOS branch performs spatial-temporal matching with the memory bank. We employ the well-learned object queries from IS branch to inject instance-specific information into the query key, with which the instance-augmented matching is further performed. In addition, we introduce a multi-path fusion block to effectively combine the memory readout with multi-scale features from the instance segmentation decoder, which incorporates high-resolution instance-aware features to produce final segmentation results. Our method achieves state-of-the-art performance on DAVIS 2016/2017 val (92.6% and 87.1%), DAVIS 2017 test-dev (82.8%), and YouTube-VOS 2018/2019 val (86.3% and 86.3%), outperforming alternative methods by clear margins. Dongdong Chen 0001, Zuxuan Wu, Chong Luo 0001, Chuanxin Tang, Xiyang Dai, Yujia Xie, Lu Yuan 0001, Yu-Gang Jiang 0001 |
CVPR | 8 |
| 2023 | Improving Commonsense in Vision-Language Models via Knowledge Graph RiddlesabstractThis paper focuses on analyzing and improving the commonsense ability of recent popular vision-language (VL) models. Despite the great success, we observe that existing VL-models still lack commonsense knowledge/reasoning ability (e.g., “Lemons are sour”), which is a vital component towards artificial general intelligence. Through our analysis, we find one important reason is that existing large-scale VL datasets do not contain much commonsense knowledge, which motivates us to improve the commonsense of VL-models from the data perspective. Rather than collecting a new VL training dataset, we propose a more scalable strategy, i.e., “Data Augmentation with kNowledge graph linearization for CommonsensE capability” (DANCE). It can be viewed as one type of data augmentation technique, which can inject commonsense knowledge into existing VL datasets on the fly during training. More specifically, we leverage the commonsense knowledge graph (e.g., ConceptNet) and create variants of text description in VL datasets via bidirectional sub-graph sequentialization. For better commonsense evaluation, we further propose the first retrieval-based commonsense diagnostic benchmark. By conducting extensive experiments on some representative VL-models, we demonstrate that our DANCE technique is able to significantly improve the commonsense ability while maintaining the performance on vanilla retrieval tasks. The code and data are available at https://github.com/pleaseconnectwifi/DANCE. Shuquan Ye, Yujia Xie, Dongdong Chen 0001, Yichong Xu, Lu Yuan 0001, Chenguang Zhu 0001, Jing Liao 0001 |
CVPR | 2 |
| 2023 | ViMRT: a text-mining tool and search engine for automated virus mutation recognitionabstractMOTIVATION: Virus mutation is one of the most important research issues which plays a critical role in disease progression and has prompted substantial scientific publications. Mutation extraction from published literature has become an increasingly important task, benefiting many downstream applications such as vaccine design and drug usage. However, most existing approaches have low performances in extracting virus mutation due to both lack of precise virus mutation information and their development based on human gene mutations. RESULTS: We developed ViMRT, a text-mining tool and search engine for automated virus mutation recognition using natural language processing. ViMRT mainly developed 8 optimized rules and 12 regular expressions based on a development dataset comprising 830 papers of 5 human severe disease-related viruses. It achieved higher performance than other tools in a test dataset (1662 papers, 99.17% in F1-score) and has been applied well to two other viruses, influenza virus and severe acute respiratory syndrome coronavirus-2 (212 papers, 96.99% in F1-score). These results indicate that ViMRT is a high-performance method for the extraction of virus mutation from the biomedical literature. Besides, we present a search engine for researchers to quickly find and accurately search virus mutation-related information including virus genes and related diseases. AVAILABILITY AND IMPLEMENTATION: ViMRT software is freely available at http://bmtongji.cn:1225/mutation/index. Yuantao Tong, Fanglin Tan, Honglian Huang, Hui Zong, Yujia Xie, Danqi Huang, Shiyang Cheng 0004, Ziyi Wei, M. James C. Crabbe, Ying Wang 0053 |
Bioinform. | 6 |
| 2023 | MADAv2: Advanced Multi-Anchor Based Active Domain Adaptation SegmentationabstractUnsupervised domain adaption has been widely adopted in tasks with scarce annotated data. Unfortunately, mapping the target-domain distribution to the source-domain unconditionally may distort the essential structural information of the target-domain data, leading to inferior performance. To address this issue, we first propose to introduce active sample selection to assist domain adaptation regarding the semantic segmentation task. By innovatively adopting multiple anchors instead of a single centroid, both source and target domains can be better characterized as multimodal distributions, in which way more complementary and informative samples are selected from the target domain. With only a little workload to manually annotate these active samples, the distortion of the target-domain distribution can be effectively alleviated, achieving a large performance gain. In addition, a powerful semi-supervised domain adaptation strategy is proposed to alleviate the long-tail distribution problem and further improve the segmentation performance. Extensive experiments are conducted on public datasets, and the results demonstrate that the proposed approach outperforms state-of-the-art methods by large margins and achieves similar performance to the fully-supervised upperbound, i.e., 71.4% mIoU on GTA5 and 71.8% mIoU on SYNTHIA. The effectiveness of each component is also verified by thorough ablation studies. Munan Ning, Donghuan Lu, Yujia Xie, Dongdong Chen 0001, Dong Wei 0004, Yefeng Zheng 0001, Yonghong Tian 0001, Shuicheng Yan, Li Yuan 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | RIP Analysis for $\ell _{1}/\ell _{p}$ ($p> 1$) Minimization MethodabstractRecently, non-convex and non-linear metrics have been introduced in compressed sensing to promote sparsity. This letter proposes an extension of the previously proposed$\ell _{1}/\ell _{2}$minimization method for sparse recovery using the$\ell _{1}/\ell _{p}$minimization method with$p\gt 1$. We establish sufficient conditions for the$\ell _{1}/\ell _{p}$minimization to recover sparse signals under the restricted isometry property (RIP). Additionally, we develop an effective algorithm to solve the$\ell _{1}/\ell _{p}$minimization problem. Experiments show the proposed method is comparable to state-of-the-art methods for sparse signal recovery. Yujia Xie, Xinhua Su, Huanmin Ge |
IEEE Signal Process. Lett. | 1 |
| 2022 | REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question AnsweringabstractThis paper revisits visual representation in knowledge-based visual question answering (VQA) and demonstrates that using regional information in a better way can significantly improve the performance. While visual representation is extensively studied in traditional VQA, it is under-explored in knowledge-based VQA even though these two tasks share the common spirit, i.e., rely on visual input to answer the question. Specifically, we observe in most state-of-the-art knowledge-based VQA methods: 1) visual features are extracted either from the whole image or in a sliding window manner for retrieving knowledge, and the important relationship within/among object regions is neglected; 2) visual features are not well utilized in the final answering model, which is counter-intuitive to some extent. Based on these observations, we propose a new knowledge-based VQA method REVIVE, which tries to utilize the explicit information of object regions not only in the knowledge retrieval stage but also in the answering model. The key motivation is that object regions and inherent relationship are important for knowledge-based VQA. We perform extensive experiments on the standard OK-VQA dataset and achieve new state-of the-art performance, i.e., 58.0 accuracy, surpassing previous state-of-the-art method by a large margin (+3.6%). We also conduct detailed analysis and show the necessity of regional information in different framework components for knowledge-based VQA. Code is publicly available at https://github.com/yzleroy/REVIVE. Yuanze Lin, Yujia Xie, Dongdong Chen 0001, Yichong Xu, Chenguang Zhu 0001, Lu Yuan 0001 |
NeurIPS | 2 |
| 2022 | K-LITE: Learning Transferable Visual Models with External KnowledgeabstractThe new generation of state-of-the-art computer vision systems are trained from natural language supervision, ranging from simple object category names to descriptive captions. This form of supervision ensures high generality and usability of the learned visual models, based on the broad concept coverage achieved through large-scale data collection process. Alternatively, we argue that learning with external knowledge about images is a promising way which leverages a much more structured source of supervision and offers sample efficiency. In this paper, we propose K-LITE (Knowledge-augmented Language-Image Training and Evaluation), a simple strategy to leverage external knowledge for building transferable visual systems: In training, it enriches entities in natural language with WordNet and Wiktionary knowledge, leading to an efficient and scalable approach to learning image representations that uses knowledge about the visual concepts; In evaluation, the natural language is also augmented with external knowledge and then used to reference learned visual concepts (or describe new ones) to enable zero-shot and few-shot transfer of the pre-trained models. We study the performance of K-LITE on two important computer vision problems, image classification and object detection, benchmarking on 20 and 13 different existing datasets, respectively. The proposed knowledge-augmented models show significant improvement in transfer learning performance over existing methods. Our code is released at https://github.com/microsoft/klite. Sheng Shen 0001, Chunyuan Li, Xiaowei Hu 0006, Yujia Xie, Pengchuan Zhang, Zhe Gan, Lu Yuan 0001, Ce Liu 0001, Kurt Keutzer, Trevor Darrell, Anna Rohrbach, Jianfeng Gao 0001 |
NeurIPS | 4 |
| 2022 | OmniVL: One Foundation Model for Image-Language and Video-Language TasksabstractThis paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture. It adopts a unified transformer-based visual encoder for both image and video inputs, and thus can perform joint image-language and video-language pretraining. We demonstrate, for the first time, such a paradigm benefits both image and video tasks, as opposed to the conventional one-directional transfer (e.g., use image-language to help video-language). To this end, we propose a \emph{decoupled} joint pretraining of image-language and video-language to effectively decompose the vision-language modeling into spatial and temporal dimensions and obtain performance boost on both image and video tasks. Moreover, we introduce a novel unified vision-language contrastive (UniVLC) loss to leverage image-text, video-text, image-label (e.g., image classification), video-label (e.g., video action recognition) data together, so that both supervised and noisily supervised pretraining data are utilized as much as possible. Without incurring extra task-specific adaptors, OmniVL can simultaneously support visual only tasks (e.g., image classification, video action recognition), cross-modal alignment tasks (e.g., image/video-text retrieval), and multi-modal understanding and generation tasks (e.g., image/video question answering, captioning). We evaluate OmniVL on a wide range of downstream tasks and achieve state-of-the-art or competitive results with similar model size and data scale. Dongdong Chen 0001, Zuxuan Wu, Chong Luo 0001, Luowei Zhou, Yujia Xie, Ce Liu 0001, Yu-Gang Jiang 0001, Lu Yuan 0001 |
NeurIPS | 7 |
| 2022 | Visual Clues: Bridging Vision and Language Foundations for Image Paragraph CaptioningabstractPeople say, "A picture is worth a thousand words". Then how can we get the rich information out of the image? We argue that by using visual clues to bridge large pretrained vision foundation models and language models, we can do so without any extra cross-modal training. Thanks to the strong zero-shot capability of foundation models, we start by constructing a rich semantic representation of the image (e.g., image tags, object attributes / locations, captions) as a structured textual prompt, called visual clues, using a vision foundation model. Based on visual clues, we use large language model to produce a series of comprehensive descriptions for the visual content, which is then verified by the vision model again to select the candidate that aligns best with the image. We evaluate the quality of generated descriptions by quantitative and qualitative measurement. The results demonstrate the effectiveness of such a structured semantic representation. Yujia Xie, Luowei Zhou, Xiyang Dai, Lu Yuan 0001, Nguyen Bach, Ce Liu 0001, Michael Zeng 0001 |
NeurIPS | 1 |
| 2021 | A Hypergradient Approach to Robust Regression without Correspondence
Yujia Xie, Yixiu Mao, Simiao Zuo, Hongteng Xu, Xiaojing Ye, Tuo Zhao, Hongyuan Zha |
ICLR | 1 |
| 2021 | Active Image Synthesis for Efficient LabelingabstractThe great success achieved by deep neural networks attracts increasing attention from the manufacturing and healthcare communities. However, the limited availability of data and high costs of data collection are the major challenges for the applications in those fields. We propose in this work AISEL, an active image synthesis method for efficient labeling, to improve the performance of the small-data learning tasks. Specifically, a complementary AISEL dataset is generated, with labels actively acquired via a physics-based method to incorporate underlining physical knowledge at hand. An important component of our AISEL method is the bidirectional generative invertible network (GIN), which can extract interpretable features from the training images and generate physically meaningful virtual images. Our AISEL method then efficiently samples virtual images not only further exploits the uncertain regions but also explores the entire image space. We then discuss the interpretability of GIN both theoretically and experimentally, demonstrating clear visual improvements over the benchmarks. Finally, we demonstrate the effectiveness of our AISEL framework on aortic stenosis application, in which our method lowers the labeling cost by 90 percent while achieving a 15 percent improvement in prediction accuracy. Jialei Chen 0002, Yujia Xie, Kan Wang 0001, Chuck Zhang, Mani A. Vannan, Ben Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Differentiable Top-k with Optimal TransportabstractFinding the k largest or smallest elements from a collection of scores, i.e., top-k operation, is an important model component widely used in information retrieval, machine learning, and data mining. However, if the top-k operation is implemented in an algorithmic way, e.g., using bubble algorithm, the resulted model cannot be trained in an end-to-end way using prevalent gradient descent algorithms. This is because these implementations typically involve swapping indices, whose gradient cannot be computed. Moreover, the corresponding mapping from the input scores to the indicator vector of whether this element belongs to the top-k set is essentially discontinuous. To address the issue, we propose a smoothed approximation, namely SOFT (Scalable Optimal transport-based diFferenTiable) top-k operator. Specifically, our SOFT top-k operator approximates the output of top-k operation as the solution of an Entropic Optimal Transport (EOT) problem. The gradient of the SOFT operator can then be efficiently approximated based on the optimality conditions of EOT problem. We then apply the proposed operator to k-nearest neighbors algorithm and beam search algorithm. The numerical experiment demonstrates their achieve improved performance. Yujia Xie, Hanjun Dai, Minshuo Chen, Bo Dai 0001, Tuo Zhao, Hongyuan Zha, Wei Wei 0019, Tomas Pfister |
NeurIPS | 1 |
| 2019 | On Scalable and Efficient Computation of Large Scale Optimal TransportabstractOptimal Transport (OT) naturally arises in many machine learning applications, yet the heavy computational burden limits its wide-spread uses. To address the scalability issue, we propose an implicit generative learning-based framework called SPOT (Scalable Push-forward of Optimal Transport). Specifically, we approximate the optimal transport plan by a pushforward of a reference distribution, and cast the optimal transport problem into a minimax problem. We then can solve OT problems efficiently using primal dual stochastic gradient-type algorithms. We also show that we can recover the density of the optimal transport plan using neural ordinary differential equations. Numerical experiments on both synthetic and real datasets illustrate that SPOT is robust and has favorable convergence behavior. SPOT also allows us to efficiently sample from the optimal transport plan, which benefits downstream applications such as domain adaptation. Yujia Xie, Minshuo Chen, Haoming Jiang, Tuo Zhao, Hongyuan Zha |
ICML | 1 |
| 2019 | Meta Learning with Relational Information for Short SequencesabstractThis paper proposes a new meta-learning method -- named HARMLESS (HAwkes Relational Meta Learning method for Short Sequences) for learning heterogeneous point process models from a collection of short event sequence data along with a relational network. Specifically, we propose a hierarchical Bayesian mixture Hawkes process model, which naturally incorporates the relational information among sequences into point process modeling. Compared with existing methods, our model can capture the underlying mixed-community patterns of the relational network, which simultaneously encourages knowledge sharing among sequences and facilitates adaptively learning for each individual sequence. We further propose an efficient stochastic variational meta-EM algorithm, which can scale to large problems. Numerical experiments on both synthetic and real data show that HARMLESS outperforms existing methods in terms of predicting the future events. Yujia Xie, Haoming Jiang, Tuo Zhao, Hongyuan Zha |
NeurIPS | 1 |
| 2019 | A Fast Proximal Point Method for Computing Exact Wasserstein Distance
Yujia Xie, Xiangfeng Wang 0001, Hongyuan Zha |
UAI | 1 |
| 2018 | Generative Invertible Networks (GIN): Pathophysiology-Interpretable Feature Mapping and Virtual Patient Generation
Jialei Chen 0002, Yujia Xie, Kan Wang 0001, Geet Lahoti, Chuck Zhang, Mani A. Vannan, Ben Wang 0001 |
MICCAI (1) | 2 |