Di You

dblp:83/8652 · DBLP profile ↗
← Back
21ranked-venue papers
11as first author
10since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Context-Aware Diffusion-based Sequential Recommendation
abstract
Sequential recommendation aims to recommend the next item that matches a user’s interest, based on the sequence of items he/she interacted with before. Although effective, existing work suffers from the following limitations: (1) Existing diffusion-based recommendation methods have undertaken tailored refinements to the diffusion process without considering the difference between recommendation and other tasks, leading to the ignorance of the user’s personalized preferences; (2) Self-supervised contrastive learning, widely used to mitigate the data sparsity issue in sequential recommendation, typically employs random augmentation to create multiple views of user sequences. However, random augmentation can disrupt the semantic integrity and interest patterns within the sequence, resulting in semantically divergent augmented views that may misrepresent user preferences. To address these challenges, we propose the Context-Aware Diffusion-based Sequential Recommendation (CADSR) model, which leverages context information to generate more semantically consistent positive samples during contrastive learning. This ensures that the model captures both user preferences and their evolution more accurately. Extensive experiments on four public benchmark datasets show that CADSR outperforms 11 state-of-the-art baselines, achieving an average improvement of 10.94% in Recall@10 and 10.54% in NDCG@10 over the best baseline. Source code is available at https://github.com/queenjocey/CADSR.
Di You, Kyumin Lee
IEEE Big Data1
2024 WiMANS: A Benchmark Dataset for WiFi-Based Multi-user Activity Sensing
Shuokang Huang, Kaihan Li, Di You, Yichong Chen, Arvin Lin, Julie A. McCann
ECCV (42)3
2024 Alleviating Confounding Effects with Contrastive Learning in Recommendation
Di You, Kyumin Lee
ECIR (2)1
2024 CommIN: Semantic Image Communications as an Inverse Problem with INN-Guided Diffusion Models
abstract
Joint source-channel coding schemes based on deep neural networks (DeepJSCC) have recently achieved remarkable performance for wireless image transmission. However, these methods usually focus only on the distortion of the reconstructed signal at the receiver side with respect to the source at the transmitter side, rather than the perceptual quality of the reconstruction which carries more semantic information. As a result, severe perceptual distortion can be introduced under extreme conditions such as low bandwidth and low signal-to-noise ratio. In this work, we propose CommIN, which views the recovery of high-quality source images from degraded reconstructions as an inverse problem. To address this, CommIN combines Invertible Neural Networks (INN) with diffusion models, aiming for superior perceptual quality. Through experiments, we show that our CommIN significantly improves the perceptual quality compared to DeepJSCC under extreme conditions and outperforms other inverse problem approaches used in DeepJSCC.
Jiakang Chen, Di You, Deniz Gündüz, Pier Luigi Dragotti
ICASSP2
2023 End-To-End Phase Retrieval from Single-Shot Fringe Image for 3D Face Reconstruction
Zhisheng You, Jiangping Zhu, Di You, Peng Cheng 0006
ICIG (3)4
2023 INDigo: An INN-Guided Probabilistic Diffusion Algorithm for Inverse Problems
abstract
Recently it has been shown that using diffusion models for inverse problems can lead to remarkable results. However, these approaches require a closed-form expression of the degradation model and can not support complex degradations. To overcome this limitation, we propose a method (INDigo) that combines invertible neural networks (INN) and diffusion models for general inverse problems. Specifically, we train the forward process of INN to simulate an arbitrary degradation process and use the inverse as a reconstruction process. During the diffusion sampling process, we impose an additional data-consistency step that minimizes the distance between the intermediate result and the INN-optimized result at every iteration, where the INN-optimized image is composed of the coarse information given by the observed degraded image and the details generated by the diffusion process. With the help of INN, our algorithm effectively estimates the details lost in the degradation process and is no longer limited by the requirement of knowing the closed-form expression of the degradation model. Experiments demonstrate that our algorithm obtains competitive results compared with recently leading methods both quantitatively and visually. Moreover, our algorithm performs well on more complex degradation models and real-world low-quality images.
Di You, Andreas Floros 0002, Pier Luigi Dragotti
MMSP1
2022 Multi-Behavior Recommendation with Hyperbolic Geometry
abstract
Even though users interacted diversely on items (e.g., click, add-to-cart, and buy), traditional recommendations were mostly built using only the user-item interaction data on the target behavior (e.g., buy), making them suffer from the severe data sparsity issue. To alleviate the problem, recent works on multi-behavior recommendation incorporated multiple types of user-item interactions such as click, add-to-cart, and buy. However, the latest approaches are still limited by overlooking early-stage interactions, and have limited expressiveness of Euclidean geometry. To overcome these issues, in this paper, we propose a Multi-behavior Hyperbolic Graph Recommender (MB-HGR) with two novel aspects. First, it uses multiple heterogeneous graphs to learn multiple user behavior types, where each heterogeneous graph represents a user-item interaction type. This will help not only alleviate the serious data sparsity problem, but also allow the model to explicitly weight different behavior types and prevent information loss. Second, it leverages the expressiveness of the hyperbolic geometry over Euclidean geometry, where exponential growth of distances in the hyperbolic geometry matches the exponential growth of nodes in the hierarchical structures and learns better users/items representations. Experimental results on two public benchmark datasets show that on average our proposed model achieves a significant improvement of 28.32% at Recall@10 and 30.14% at NDCG@10 over the best baseline.
Di You, Thanh Tran 0005, Kyumin Lee
IEEE Big Data1
2021 ISTA-NET++: Flexible Deep Unfolding Network for Compressive Sensing
abstract
While deep neural networks have achieved impressive success in image compressive sensing (CS), most of them lack flexibility when dealing with multi-ratio tasks and multi-scene images in practical applications. To tackle these challenges, we propose a novel end-to-end flexible ISTA-unfolding deep network, dubbed ISTA-Net++, with superior performance and strong flexibility. Specifically, by developing a dynamic unfolding strategy, our model enjoys the adaptability of handling CS problems with different ratios, i.e., multi-ratio tasks, through a single model. A cross-block strategy is further utilized to reduce blocking artifacts and enhance the CS recovery quality. Furthermore, we adopt a balanced dataset for training, which brings more robustness when reconstructing images of multiple scenes. Extensive experiments on four datasets show that ISTA-Net++achieves state-of-the-art results in terms of both quantitative metrics and visual quality. Considering its flexibility, effectiveness and practicability, our model is expected to serve as a suitable baseline in future CS research. The source code is available on https://github.com/jianzhangcs/ISTA-Netpp.
Di You, Jingfen Xie, Jian Zhang 0018
ICME1
2021 AlignTransformer: Hierarchical Alignment of Visual Regions and Disease Tags for Medical Report Generation
Di You, Shen Ge, Xiaoxia Xie, Xian Wu 0001
MICCAI (3)1
2021 COAST: COntrollable Arbitrary-Sampling NeTwork for Compressive Sensing
abstract
Recent deep network-based compressive sensing (CS) methods have achieved great success. However, most of them regard different sampling matrices as different independent tasks and need to train a specific model for each target sampling matrix. Such practices give rise to inefficiency in computing and suffer from poor generalization ability. In this paper, we propose a novel COntrollable Arbitrary-Sampling neTwork, dubbed COAST, to solve CS problems of arbitrary-sampling matrices (including unseen sampling matrices) with one single model. Under the optimization-inspired deep unfolding framework, our COAST exhibits good interpretability. In COAST, a random projection augmentation (RPA) strategy is proposed to promote the training diversity in the sampling space to enable arbitrary sampling, and a controllable proximal mapping module (CPMM) and a plug-and-play deblocking (PnP-D) strategy are further developed to dynamically modulate the network features and effectively eliminate the blocking artifacts, respectively. Extensive experiments on widely used benchmark datasets demonstrate that our proposed COAST is not only able to handle arbitrary sampling matrices with one single model but also to achieve state-of-the-art performance with fast speed.
Di You, Jian Zhang 0018, Jingfen Xie, Bin Chen 0006, Siwei Ma 0001
IEEE Trans. Image Process.1
2020 Quaternion-Based Self-Attentive Long Short-term User Preference Encoding for Recommendation
abstract
Quaternion space has brought several benefits over the traditional Euclidean space: Quaternions (i) consist of a real and three imaginary components, encouraging richer representations; (ii) utilize Hamilton product which better encodes the inter-latent interactions across multiple Quaternion components; and (iii) result in a model with smaller degrees of freedom and less prone to overfitting. Unfortunately, most of the current recommender systems rely on real-valued representations in Euclidean space to model either user's long-term or short-term interests. In this paper, we fully utilize Quaternion space to model both user's long-term and short-term preferences. We first propose a QUaternion-based self-Attentive Long term user Encoding (QUALE) to study the user's long-term intents. Then, we propose a QUaternion-based self-Attentive Short term user Encoding (QUASE) to learn the user's short-term interests. To enhance our models' capability, we propose to fuse QUALE and QUASE into one model, namely QUALSE, by using a Quaternion-based gating mechanism. We further develop Quaternion-based Adversarial learning along with the Bayesian Personalized Ranking (QABPR) to improve our model's robustness. Extensive experiments on six real-world datasets show that our fused QUALSE model outperformed 11 state-of-the-art baselines, improving 8.43% at [email protected] and 10.27% at [email protected] on average compared with the best baseline.
Thanh Tran 0005, Di You, Kyumin Lee
CIKM2
2019 Bridging the Gap between Training and Inference for Neural Machine Translation
abstract
Neural Machine Translation (NMT) generates target words sequentially in the way of predicting the next word conditioned on the context words.At training time, it predicts with the ground truth words as context while at inference it has to generate the entire sequence from scratch.This discrepancy of the fed context leads to error accumulation among the way.Furthermore, word-level training requires strict matching between the generated sequence and the ground truth sequence which leads to overcorrection over different but reasonable translations.In this paper, we address these issues by sampling context words not only from the ground truth sequence but also from the predicted sequence by the model during training, where the predicted sequence is selected with a sentence-level optimum.Experiment results on Chinese→English and WMT'14 English→German translation tasks demonstrate that our approach can achieve significant improvements on multiple datasets.
Wen Zhang 0009, Yang Feng 0004, Fandong Meng, Di You, Qun Liu 0001
ACL (1)4
2019 Detecting Fake News Articles
abstract
Fake news has been generated and widely spread although journalists and researchers created fact-checking websites (e.g., Snopes and PolitiFact) and analyzed characteristics of fake news. To fill this gap, in this paper we focus on developing machine learning models based on only text information in news articles toward automatically detecting fake news. In particular, we proposed a framework which extracts 134 features and builds traditional known machine learning models like Random Forest and XGBoost. We also propose a deep learning based model (LSTM with self-attention mechanism) to see which one performs better in the fake news article detection in both political news and celebrity news domains. In the experiments, we compare our models against 7 baselines. The results show that our XGBoost model improved 16.4% and 13.1% over the best baseline in terms of accuracy in both political news articles and celebrity news articles, respectively.
Glenna Tremblay-Taylor, Guanyi Mou, Di You, Kyumin Lee
IEEE BigData4
2019 Attributed Multi-Relational Attention Network for Fact-checking URL Recommendation
abstract
To combat fake news, researchers mostly focused on detecting fake news and journalists built and maintained fact-checking sites (e.g., Snopes.com and Politifact.com). However, fake news dissemination has been greatly promoted via social media sites, and these fact-checking sites have not been fully utilized. To overcome these problems and complement existing methods against fake news, in this paper we propose a deep-learning based fact-checking URL recommender system to mitigate impact of fake news in social media sites such as Twitter and Facebook. In particular, our proposed framework consists of a multi-relational attentive module and a heterogeneous graph attention network to learn complex/semantic relationship between user-URL pairs, user-user pairs, and URL-URL pairs. Extensive experiments on a real-world dataset show that our proposed framework outperforms eight state-of-the-art recommendation models, achieving at least 3$\sim$5.3% improvement. Our source code and dataset are available at \urlhttps://web.cs.wpi.edu/~kmlee/data.html .
Di You, Nguyen Vo, Kyumin Lee
CIKM1
2019 Neural Machine Translation with Bilingual History Involved Attention
Haiyang Xue, Yang Feng 0004, Di You, Wen Zhang 0009
NLPCC (2)3
2014 Multiobjective Optimization for Model Selection in Kernel Methods in Regression
abstract
Regression plays a major role in many scientific and engineering problems. The goal of regression is to learn the unknown underlying function from a set of sample vectors with known outcomes. In recent years, kernel methods in regression have facilitated the estimation of nonlinear functions. However, two major (interconnected) problems remain open. The first problem is given by the bias-versus-variance tradeoff. If the model used to estimate the underlying function is too flexible (i.e., high model complexity), the variance will be very large. If the model is fixed (i.e., low complexity), the bias will be large. The second problem is to define an approach for selecting the appropriate parameters of the kernel function. To address these two problems, this paper derives a new smoothing kernel criterion, which measures the roughness of the estimated function as a measure of model complexity. Then, we use multiobjective optimization to derive a criterion for selecting the parameters of that kernel. The goal of this criterion is to find a tradeoff between the bias and the variance of the learned function. That is, the goal is to increase the model fit while keeping the model complexity in check. We provide extensive experimental evaluations using a variety of problems in machine learning, pattern recognition, and computer vision. The results demonstrate that the proposed approach yields smaller estimation errors as compared with methods in the state of the art.
Di You, Carlos F. Benitez-Quiroz, Aleix Martinez
IEEE Trans. Neural Networks Learn. Syst.1
2013 Towards zero-shot learning for human activity recognition using semantic attribute sequence model
abstract
Understanding human activities is important for user-centric and context-aware applications. Previous studies showed promising results using various machine learning algorithms. However, most existing methods can only recognize the activities that were previously seen in the training data. In this paper, we present a new zero-shot learning framework for human activity recognition that can recognize an unseen new activity even when there are no training samples of that activity in the dataset. We propose a semantic attribute sequence model that takes into account both the hierarchical and sequential nature of activity data. Evaluation on datasets in two activity domains show that the proposed zero-shot learning approach achieves 70-75% precision and recall recognizing unseen new activities, and outperforms supervised learning with limited labeled data for the new classes.
Heng-Tze Cheng, Martin L. Griss, Di You
UbiComp5
2013 NuActiv: recognizing unseen new activities using semantic attribute-based learning
abstract
We study the problem of how to recognize a new human activity when we have never seen any training example of that activity before. Recognizing human activities is an essential element for user-centric and context-aware applications. Previous studies showed promising results using various machine learning algorithms. However, most existing methods can only recognize the activities that were previously seen in the training data. A previously unseen activity class cannot be recognized if there were no training samples in the dataset. Even if all of the activities can be enumerated in advance, labeled samples are often time consuming and expensive to get, as they require huge effort from human annotators or experts. In this paper, we present NuActiv, an activity recognition system that can recognize a human activity even when there are no training data for that activity class. Firstly, we designed a new representation of activities using semantic attributes, where each attribute is a human readable term that describes a basic element or an inherent characteristic of an activity. Secondly, based on this representation, a two-layer zero-shot learning algorithm is developed for activity recognition. Finally, to reinforce recognition accuracy using minimal user feedback, we developed an active learning algorithm for activity recognition. Our approach is evaluated on two datasets, including a 10-exercise-activity dataset we collected, and a public dataset of 34 daily life activities. Experimental results show that using semantic attribute-based learning, NuActiv can generalize knowledge to recognize unseen new activities. Our approach achieved up to 79% accuracy in unseen activity recognition.
Heng-Tze Cheng, Feng-Tso Sun, Martin L. Griss, Di You
MobiSys6
2012 Representative Multiple Kernel Learning for Classification in Hyperspectral Imagery
abstract
Recently, multiple kernel learning (MKL) methods have been developed to improve the flexibility of kernel-based learning machine. The MKL methods generally focus on determining key kernels to be preserved and their significance in optimal kernel combination. Unfortunately, computational demand of finding the optimal combination is prohibitive when the number of training samples and kernels increase rapidly, particularly for hyperspectral remote sensing data. In this paper, we address the MKL for classification in hyperspectral images by extracting the most variation from the space spanned by multiple kernels and propose a representative MKL (RMKL) algorithm. The core idea embedded in the algorithm is to determine the kernels to be preserved and their weights according to statistical significance instead of time-consuming search for optimal kernel combination. The noticeable merits of RMKL consist that it greatly reduces the computational load for searching optimal combination of basis kernels and has no limitation from strict selection of basis kernels like most MKL algorithms do; meanwhile, RMKL keeps excellent properties of MKL in terms of both good classification accuracy and interpretability. Experiments are conducted on different real hyperspectral data, and the corresponding experimental results show that RMKL algorithm provides the best performances to date among several the state-of-the-art algorithms while demonstrating satisfactory computational efficiency.
Yanfeng Gu, Di You, Yuhang Zhang 0002, Shizhe Wang, Ye Zhang 0008
IEEE Trans. Geosci. Remote. Sens.3
2011 Kernel Optimization in Discriminant Analysis
abstract
Kernel mapping is one of the most used approaches to intrinsically derive nonlinear classifiers. The idea is to use a kernel function which maps the original nonlinearly separable problem to a space of intrinsically larger dimensionality where the classes are linearly separable. A major problem in the design of kernel methods is to find the kernel parameters that make the problem linear in the mapped representation. This paper derives the first criterion that specifically aims to find a kernel representation where the Bayes classifier becomes linear. We illustrate how this result can be successfully applied in several kernel discriminant analysis algorithms. Experimental results, using a large number of databases and classifiers, demonstrate the utility of the proposed approach. The paper also shows (theoretically and experimentally) that a kernel version of Subclass Discriminant Analysis yields the highest recognition rates.
Di You, Onur C. Hamsici, Aleix Martinez
IEEE Trans. Pattern Anal. Mach. Intell.1
2010 Bayes optimal kernel discriminant analysis
abstract
Kernel methods provide an efficient mechanism to derive nonlinear algorithms. In classification problems as well as in feature extraction, kernel-based approaches map the originally nonlinearly separable data into a space of intrinsically much higher dimensionality where the data is linearly separable and can be readily classified with existing and efficient linear methods. For a given kernel function, the main challenge is to determine the parameters of the kernel which maps the original nonlinear problem to a linear one. This paper derives a Bayes optimal criterion for the selection of the kernel parameters in discriminant analysis. Our criterion selects the kernel parameters that maximize the (Bayes) classification accuracy in the kernel space. We also show how we can use the same criterion to do subclass selection in the kernel space for problems with multimodal class distributions. Extensive experimental evaluation demonstrates the superiority of the proposed criterion over the state of the art.
Di You, Aleix Martinez
CVPR1