Daniel Andrade

dblp:32/9125 · DBLP profile ↗
← Back
19ranked-venue papers
14as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 12 first-author · 3 since 2021Security and privacy · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Critical Factors for a Reliable AI in Tutoring Systems on Accuracy, Effectiveness, and Responsibility
abstract
In recent years, there has been a surge in the development and use of artificial intelligence (AI) systems in various fields, including education. One such application is the AI-based tutoring system, which can provide personalized learning experiences to students. These systems leverage advanced algorithms to analyze student performance, identify knowledge gaps, and deliver targeted feedback and guidance. One of the significant challenges educators and researchers face in the context of AI-based tutoring systems is the lack of reliability in the systems. The accuracy, effectiveness, and responsibility of AI systems are critical factors that determine their reliability. Accuracy challenges for AI algorithms in tutoring systems include accurately modeling individual learner profiles, providing tailored content that aligns with each student's pace and understanding, addressing diverse learning strategies, and ensuring the feedback is specific and actionable. Overcoming data sparsity and ensuring algorithmic transparency and fairness are also significant hurdles. Challenges in ensuring effectiveness include developing algorithms that accurately adapt to individual learning needs, processing natural language effectively, maintaining engagement, and providing contextually relevant feedback. Responsibility challenges include ensuring data privacy and security, preventing algorithm biases affecting learning outcomes, and maintaining ethical standards in AI interactions. Balancing automation with human oversight to support diverse learning needs without compromising educational integrity is also crucial. Considering these challenges, this study discusses the critical factors for reliable AI in tutoring systems from perspectives of accuracy, effectiveness, and responsibility. The research has a descriptive character and qualitative analysis, applying the systematic literature review (SLR) method. The analysis of 43 studies in the last five years made it possible to find some interesting results. In summary, the accuracy of AI in Tutoring Systems is impacted by data quality and preprocessing, choice of appropriate metrics, advanced learning techniques, dataset diversity, and correct selection of hyperparameters. As for the factors that influence the effectiveness of these systems, the personalization of learning and the ability of the systems to adapt to individual needs through personalized pedagogical interventions stand out, using techniques such as recurrent neural networks to predict the quality of interactions. However, challenges related to understanding learning emotions reinforce the complexity of building effective models based on emotion induction. Concerning the responsible use of AI in tutoring systems, it is crucial to consider the privacy and security of student data, adopt collaborative human-machine approaches, and align the use of AI with institutional governance, promoting a safe and ethical learning environment. As a main contribution, we highlight an enlightening discussion of the critical factors for a reliable AI in the context of tutoring systems, identifying quality studies on the subject to support researchers.
Davi Maia, Simone C. dos Santos, Luis Gabriel Lima, Vinícius Luiz Franca, Alexsandro Henrique Lima, Daniel Andrade
FIE6
2023 I Can't Escape Myself: Cloud Inter-Processor Attestation and Sealing using Intel SGX
abstract
SGX enclaves protect the code and data within from untrusted software, but do not retain their state when destroyed. Sealing allows preserving this data for future use by storing it outside the enclave boundary: the data is encrypted with a fresh secret key, and this key is bound to the processor that sealed the data and, either to the enclave measurement, or the public key of the enclave author. However, in a cloud environment, customers do not choose in which processor their code executes: the enclave that seals some data (for backup or future use) may be destroyed and later instantiated on a different processor or migrated to another processor. In those cases, the new processor would not be able to unseal the data, since the secret key is bound to the sealing processor.This paper presents Inter-Processor Attestation and Sealing (IPAS), a novel sealing mechanism for cloud environments and Intel SGX. In IPAS, sealed data is no longer bound to the sealing processor, but only to the enclave measurement or the public key of the enclave author, thus enabling other processors to unseal the sealed data. This is achieved without exporting the sealing secret key outside the enclave and without trusting third parties.
Daniel Andrade, João Nuno Silva, Miguel Correia 0001
PRDC1
2023 TrustGlass: Human-Computer Trusted Paths with Augmented Reality Smart Glasses
abstract
Humans constantly interact with computing devices. Many times, these interactions involve sensitive information and/or sensitive commands, leading to the need for trusted paths between humans and computers. For computer-to-computer interactions, secure communication is not an issue thanks to cryptographic protocols such as Transport Layer Security (TLS) and the IP Security Protocol (IPSec). However, the same does not yet apply to interactions between humans and computers due to the inability of humans to execute non-trivial cryptographic algorithms. This results in interactions that are vulnerable to shoulder surfing, man-in-the-browser malware, and web page spoofing, among other attacks. This consequently puts at risk the use of sensitive services in uncontrolled environments.We present TrustGlass, a scheme that uses augmented reality smart glasses to solve this security problem. TrustGlass uses such glasses to extend human capabilities, to execute the cryptographic algorithms needed to establish a trusted path between the user and the trusted service, which can be executed in a Trusted Execution Environment (TEE). With this approach, we can guarantee that only the user equipped with the glasses is capable of interacting with the service. We implemented TrustGlass using commercial smart glasses and show experimental performance and usability results.
Hélio Borges, Daniel Andrade, João Nuno Silva, Miguel Correia 0001
TrustCom2
2023 Robust Gaussian process regression with the trimmed marginal likelihood
abstract
Accurate outlier detection is not only a necessary preprocessing step, but can itself give important insights into the data. However, especially, for non-linear regression the detection of outliers is non-trivial, and actually ambiguous. We propose a new method that identifies outliers by finding a subset of data points T such that the marginal likelihood of all remaining data points S is maximized. Though the idea is more general, it is particular appealing for Gaussian processes regression, where the marginal likelihood has an analytic solution. While maximizing the marginal likelihood for hyper-parameter optimization is a well established non-convex optimization problem, optimizing the set of data points S is not. Indeed, even a greedy approximation is computationally challenging due to the high cost of evaluating the marginal likelihood. As a remedy, we propose an efficient projected gradient descent method with provable convergence guarantees. Moreover, we also establish the breakdown point when jointly optimizing hyper-parameters and S. For various datasets and types of outliers, our experiments demonstrate that the proposed method can improve outlier detection and robustness when compared with several popular alternatives like the student-t likelihood.
Daniel Andrade, Akiko Takeda
UAI1
2022 Anonymous Trusted Data Relocation for TEEs
Vasco Guita, Daniel Andrade, João Nuno Silva, Miguel Correia 0001
SEC2
2021 Adaptive covariate acquisition for minimizing total cost of classification
abstract
Abstract In some applications, acquiring covariates comes at a cost which is not negligible. For example in the medical domain, in order to classify whether a patient has diabetes or not, measuring glucose tolerance can be expensive. Assuming that the cost of each covariate, and the cost of misclassification can be specified by the user, our goal is to minimize the (expected) total cost of classification, i.e. the cost of misclassification plus the cost of the acquired covariates. We formalize this optimization goal using the (conditional) Bayes risk and describe the optimal solution using a recursive procedure. Since the procedure is computationally infeasible, we consequently introduce two assumptions: (1) the optimal classifier can be represented by a generalized additive model, (2) the optimal sets of covariates are limited to a sequence of sets of increasing size. We show that under these two assumptions, a computationally efficient solution exists. Furthermore, on several medical datasets, we show that the proposed method achieves in most situations the lowest total costs when compared to various previous methods. Finally, we weaken the requirement on the user to specify all misclassification costs by allowing the user to specify the minimally acceptable recall (target recall). Our experiments confirm that the proposed method achieves the target recall while minimizing the false discovery rate and the covariate acquisition costs better than previous methods.
Daniel Andrade, Yuzuru Okajima
Mach. Learn.1
2021 Convex covariate clustering for classification
abstract
Clustering, like covariate selection for classification, is an important step to compress and interpret the data. However, clustering of covariates is often performed independently of the classification step, which can lead to undesirable clustering results that harm interpretability and compression rate. Therefore, we propose a method that can cluster covariates while taking into account class label information of samples. We formulate the problem as a convex optimization problem which uses both, a-priori similarity information between covariates, and information from class-labeled samples. Like ordinary convex clustering [1], the proposed method offers a unique global minima making it insensitive to initialization. In order to solve the convex problem, we propose a specialized alternating direction method of multipliers (ADMM), which scales up to several thousands of variables. Furthermore, in order to circumvent computationally expensive cross-validation, we propose a model selection criterion based on approximating the marginal likelihood. Experiments on synthetic and real data confirm the usefulness of the proposed clustering method and the selection criterion.
Daniel Andrade, Kenji Fukumizu, Yuzuru Okajima
Pattern Recognit. Lett.1
2020 POSTER: Detecting Suspicious Processes from Log-Data via a Bayesian Block Model
abstract
Analyzing the behavior of an attacker is critical for determining the scope of damage of a cyber attack, recovering, and fixing system vulnerabilities. However, finding all attacker's traces from log data is a laborsome task, where the performance of existing machine learning methods is still insufficient. In this work, we focus on the task of detecting all processes that were executed by the attacker. For this task, standard anomaly detection methods like Isolation Forest, perform poorly, due to many processes that are used by both the attacker and the client user. Therefore, we propose to incorporate prior knowledge about the temporal concentration of the attacker's activity. In general, we expect that an attacker is active only during a relatively small time window (block assumption), rather than being active at completely random time points. We propose a generative model that allows us to incorporate such prior knowledge effectively. Experiments on intrusion log data, shows that the proposed method achieves considerably better detection performance than a strong baseline method which also incorporates the block assumption.
Daniel Andrade, Yusuke Takahashi, Daichi Hasumi
AsiaCCS1
2020 Intelligent Epidemiological Surveillance in the Brazilian Semiarid
abstract
Right after the Chinese example in conducting COVID-19 epidemic originated in Wuhan, the readiness to detect and respond by health authorities to local (sometimes global) epidemics has become central lately. Within the idea of health 4.0, information about the individual is essential in supporting public community health policies. This paper presents a proposal for an epidemiological surveillance system applied to arboviruses. Data mining techniques and Machine Learning (ML) are used to design mathematical models for detecting epidemics enhanced by Aedes Aegypti (vector for dengue, chikungunaya, yellow fever and zica). Based on data, it is proposed an adaptive manner to reach better stability on results. A Prove of Concept (PoC) is presented for dengue epidemics detection, a common endemic disease in the semiarid region of Brazil.
Raimundo Valter, Mauro Oliveira, William Vitorino, José Neuman de Souza, Samuel Albuquerque, Luiz O. M. Andrade, Ivana Cristina de H. C. Barreto, Flávio Cardoso, Francisco G. S. da Silva, Daniel Andrade, Luzia Lucélia Saraiva Ribeiro
HealthCom10
2019 Efficient Bayes Risk Estimation for Cost-Sensitive Classification
abstract
In some real world applications, acquiring covariates for classification can be cost-intensive and should be limited as much as possible. For example, in the medical setting, a doctor cannot just perform all possible types of tests to classify whether the patient has diabetes or not. The decision of classifying or acquiring more covariates before classifying is dependent on the costs of new covariates and the expected optimal cost of misclassification (Bayes risk). However, estimating the latter is a formidable task due to the estimation of a high dimensional probability density and intractable integrals. In this work, we show that for linear classifiers this task can be considerably simplified, leading to a one dimensional integral for which we propose an efficient approximation. Experimental results on three datasets show consistent improvements over previously proposed methods for cost-sensitive classification. We also demonstrate that our proposed Bayes risk estimation procedure can benefit from additional unlabeled data which can be helpful when only small amount of labeled data is available.
Daniel Andrade, Yuzuru Okajima
AISTATS1
2018 Exploiting covariate embeddings for classification using Gaussian processes
Daniel Andrade, Akihiro Tamura, Masaaki Tsuchida
Pattern Recognit. Lett.1
2015 Cross-lingual Text Classification Using Topic-Dependent Word Probabilities
abstract
Daniel Andrade, Kunihiko Sadamasa, Akihiro Tamura, Masaaki Tsuchida. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Daniel Andrade, Kunihiko Sadamasa, Akihiro Tamura, Masaaki Tsuchida
HLT-NAACL1
2013 Synonym Acquisition Using Bilingual Comparable Corpora
Daniel Andrade, Masaaki Tsuchida, Takashi Onishi, Kai Ishikawa
IJCNLP1
2013 Chinese Informal Word Normalization: an Experimental Study
Aobo Wang, Min-Yen Kan, Daniel Andrade, Takashi Onishi, Kai Ishikawa
IJCNLP3
2013 Translation Acquisition Using Synonym Sets
Daniel Andrade, Masaaki Tsuchida, Takashi Onishi, Kai Ishikawa
HLT-NAACL1
2012 Statistical Extraction and Comparison of Pivot Words for Bilingual Lexicon Extension
abstract
Bilingual dictionaries can be automatically extended by new translations using comparable corpora. The general idea is based on the assumption that similar words have similar contexts across languages. However, previous studies have mainly focused on Indo-European languages, or use only a bag-of-words model to describe the context. Furthermore, we argue that it is helpful to extract only the statistically significant context, instead of using all context. The present approach addresses these issues in the following manner. First, based on the context of a word with an unknown translation (query word), we extract salient pivot words. Pivot words are words for which a translation is already available in a bilingual dictionary. For the extraction of salient pivot words, we use a Bayesian estimation of the point-wise mutual information to measure statistical significance. In the second step, we match these pivot words across languages to identify translation candidates for the query word. We therefore calculate a similarity score between the query word and a translation candidate using the probability that the same pivots will be extracted for both the query word and the translation candidate. The proposed method uses several context positions, namely, a bag-of-words of one sentence, and the successors, predecessors, and siblings with respect to the dependency parse tree of the sentence. In order to make these context positions comparable across Japanese and English, which are unrelated languages, we use several heuristics to adjust the dependency trees appropriately. We demonstrate that the proposed method significantly increases the accuracy of word translations, as compared to previous methods.
Daniel Andrade, Takuya Matsuzaki, Jun'ichi Tsujii
ACM Trans. Asian Lang. Inf. Process.1
2011 Effective Use of Dependency Structure for Bilingual Lexicon Creation
Daniel Andrade, Takuya Matsuzaki, Jun'ichi Tsujii
CICLing (2)1
2010 Robust Measurement and Comparison of Context Similarity for Finding Translation Pairs
Daniel Andrade, Tetsuya Nasukawa, Jun'ichi Tsujii
COLING1
2009 Lower Bound Bayesian Networks - An Efficient Inference of Lower Bounds on Probability Distributions in Bayesian Networks
Daniel Andrade, Bernhard Sick
UAI1