EDBT 2026 Demo / reviewers in the wild / expert
Issam H. Laradji
dblp:142/0043 · also Issam Hadj Laradji
· DBLP profile ↗
38ranked-venue papers
9as first author
22since 2021 · last 2025
0000-0002-9713-3269ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 11 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | StarVector: Generating Scalable Vector Graphics Code from Images and TextabstractScalable Vector Graphics (SVG) have become integral to modern image rendering applications due to their infinite scalability and versatility, especially in graphic design and web development. SVGs are essentially long strings of code that adhere to a structured syntax with validity constraints. With the rise of large language models, which excel at generating code in various languages, we aim to generate SVG code in a similar way. Our findings show that a vision-language model can be conditioned to produce valid SVG code that closely resembles input images, effectively enabling vectorization. Additionally, we harness the rich SVG syntax, encompassing all possible primitives—such as lines, paths, polygons, text, and effects like color gradients—that previous methods often missed. We briefly explain how the StarVector model operates, primarily leveraging a vision-language transformer architecture to generate SVG code. We also detail our training and inference procedures. Finally, we provide an interactive demo that allows users to input an image and generate its SVG code autoregressively, featuring real-time rendering that visually demonstrates the SVG generation process. Juan A. Rodríguez, Abhay Puri, Issam H. Laradji, Sai Rajeswar, David Vázquez 0001, Christopher Joseph Pal, Marco Pedersoli |
AAAI | 4 |
| 2025 | Fast Convergence of Softmax Policy Mirror AscentabstractNatural policy gradient (NPG) is a common policy optimization algorithm and can be viewed as mirror ascent in the space of probabilities. Recently, Vaswani et al. (2021) introduced a policy gradient method that corresponds to mirror ascent in the dual space of logits. We refine this algorithm, removing its need for a normalization across actions and analyze the resulting method (referred to as SPMA). For tabular MDPs, we prove that SPMA with a constant step-size matches the linear convergence of NPG and achieves a faster convergence than constant step-size (accelerated) softmax policy gradient. To handle large state-action spaces, we extend SPMA to use a log-linear policy parameterization. Unlike that for NPG, generalizing SPMA to the linear function approximation (FA) setting does not require compatible function approximation. Unlike MDPO, a practical generalization of NPG, SPMA with linear FA only requires solving convex softmax classification problems. We prove that SPMA achieves linear convergence to the neighbourhood of the optimal value function. We extend SPMA to handle non-linear FA and evaluate its empirical performance on the MuJoCo and Atari benchmarks. Our results demonstrate that SPMA consistently achieves similar or better performance compared to MDPO, PPO and TRPO. Reza Asad, Reza Babanezhad 0001, Issam H. Laradji, Nicolas Le Roux, Sharan Vaswani |
AISTATS | 3 |
| 2025 | StarVector: Generating Scalable Vector Graphics Code from Images and TextabstractScalable Vector Graphics (SVGs) are vital for modern image rendering due to their scalability and versatility. Previous SVG generation methods have focused on curve-based vectorization, lacking semantic understanding, often producing artifacts, and struggling with SVG primitives beyond path curves. To address these issues, we introduce StarVector, a multimodal large language model for SVG generation. It performs image vectorization by understanding image semantics and using SVG primitives for compact, precise outputs. Unlike traditional methods, StarVector works directly in the SVG code space, leveraging visual understanding to apply accurate SVG primitives. To train StarVector, we create SVG-Stack, a diverse dataset of 2M samples that enables generalization across vectorization tasks and precise use of primitives like ellipses, polygons, and text. We address challenges in SVG evaluation, showing that pixel-based metrics like MSE fail to capture the unique qualities of vector graphics. We introduce SVG-Bench, a benchmark across 10 datasets, and 3 tasks: Image-to-SVG, Text-to-SVG generation, and diagram generation. Using this setup, StarVector achieves state-of-the-art performance, producing more compact and semantically rich SVGs. Juan A. Rodríguez, Abhay Puri, Issam H. Laradji, Pau Rodríguez, Sai Rajeswar, David Vázquez 0001, Christopher Joseph Pal, Marco Pedersoli |
CVPR | 4 |
| 2025 | InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight GenerationabstractData analytics is essential for extracting valuable insights from data that can assist organizations in making effective decisions. We introduce InsightBench, a benchmark dataset with three key features. First, it consists of 100 datasets representing diverse business use cases such as finance and incident management, each accompanied by a carefully curated set of insights planted in the datasets. Second, unlike existing benchmarks focusing on answering single queries, InsightBench evaluates agents based on their ability to perform end-to-end data analytics, including formulating questions, interpreting answers, and generating a summary of insights and actionable steps. Third, we conducted comprehensive quality assurance to ensure that each dataset in the benchmark had clear goals and included relevant and meaningful questions and analysis. Furthermore, we implement a two-way evaluation mechanism using LLaMA-3 as an effective, open-source evaluator to assess agents’ ability to extract insights. We also propose AgentPoirot, our baseline data analysis agent capable of performing end-to-end data analytics. Our evaluation on InsightBench shows that AgentPoirot outperforms existing approaches (such as Pandas Agent) that focus on resolving single queries. We also compare the performance of open- and closed-source LLMs and various evaluation strategies. Overall, this benchmark serves as a testbed to motivate further development in comprehensive automated data analytics and can be accessed here: https://github.com/ServiceNow/insight-bench. Gaurav Sahu, Abhay Puri, Juan A. Rodríguez, Amirhossein Abaskohi, Mohammad Chegini, Alexandre Drouin, Perouz Taslakian, Valentina Zantedeschi, Alexandre Lacoste, David Vázquez 0001, Nicolas Chapados, Christopher Joseph Pal, Sai Rajeswar, Issam H. Laradji |
ICLR | 14 |
| 2025 | AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingabstractAligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared embedding space with the LLM while preserving semantic similarity. Existing connectors, such as multilayer perceptrons (MLPs), lack inductive bias to constrain visual features within the linguistic structure of the LLM’s embedding space, making them data-hungry and prone to cross-modal misalignment. In this work, we propose a novel vision-text alignment method, AlignVLM, that maps visual features to a weighted average of LLM text embeddings. Our approach leverages the linguistic priors encoded by the LLM to ensure that visual features are mapped to regions of the space that the LLM can effectively interpret. AlignVLM is particularly effective for document understanding tasks, where visual and textual modalities are highly correlated. Our extensive experiments show that AlignVLM achieves state-of-the-art performance compared to prior alignment methods, with larger gains on document understanding and under low-resource setups. We provide further analysis demonstrating its efficiency and robustness to noise. Ahmed Masry, Juan A. Rodríguez, Suyuchen Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu 0003, Nicolas Chapados, Yoshua Bengio, Enamul Hoque Prince, Christopher Joseph Pal, Issam H. Laradji, David Vázquez 0001, Perouz Taslakian, Spandana Gella, Sai Rajeswar |
NeurIPS | 18 |
| 2024 | WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?abstractWe study the use of large language model-based agents for interacting with software via web browsers. Unlike prior work, we focus on measuring the agents’ ability to perform tasks that span the typical daily work of knowledge workers utilizing enterprise software systems. To this end, we propose WorkArena, a remote-hosted benchmark of 33 tasks based on the widely-used ServiceNow platform. We also introduce BrowserGym, an environment for the design and evaluation of such agents, offering a rich set of actions as well as multimodal observations. Our empirical evaluation reveals that while current agents show promise on WorkArena, there remains a considerable gap towards achieving full task automation. Notably, our analysis uncovers a significant performance disparity between open and closed-source LLMs, highlighting a critical area for future exploration and development in the field. Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vázquez 0001, Nicolas Chapados, Alexandre Lacoste |
ICML | 4 |
| 2023 | PromptMix: A Class Boundary Augmentation Method for Large Language Model DistillationabstractData augmentation is a widely used technique to address the problem of text classification when there is a limited amount of training data.Recent work often tackles this problem using large language models (LLMs) like GPT3 that can generate new examples given already available ones.In this work, we propose a method to generate more helpful augmented data by utilizing the LLM's abilities to follow instructions and perform few-shot classifications.Our specific PromptMix method consists of two steps: 1) generate challenging text augmentations near class boundaries; however, generating borderline examples increases the risk of false positives in the dataset, so we 2) relabel the text augmentations using a prompting-based LLM classifier to enhance the correctness of labels in the generated data.We evaluate the proposed method in challenging 2-shot and zero-shot settings on four text classification datasets: Bank-ing77, TREC6, Subjectivity (SUBJ), and Twitter Complaints.Our experiments show that generating and, crucially, relabeling borderline examples facilitates the transfer of knowledge of a massive LLM like GPT3.5-turbo into smaller and cheaper classifiers like DistilBERT base and BERT base .Furthermore, 2-shot Prompt-Mix outperforms multiple 5-shot data augmentation methods on the four datasets. Gaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, Issam H. Laradji |
EMNLP | 4 |
| 2023 | Affinity Learning With Blind-Spot Self-Supervision for Image DenoisingabstractIn this paper, we extend the blind-spot based self-supervised denoising by using affinity learning to remove noise from affected pixels. Inspired by inpainting, we introduce a novel Mask Guided Residual Convolution (MGRConv) to learn a neighboring image pixel affinity map that gradually removes noise and refines blind-spot denoising process. We show that mask convolution plays an important role in blind-spot denoising since it is theoretically aligned with $\mathcal{J} - invariance$, which blind-spot based self-supervised denoising frameworks are built upon. The theoretical analysis further shows the motivation behind using more adaptive mask convolutions. Our MGRConv not only enables dynamic mask learning without external trainable parameters, but also preserves appropriate mask constraints by sigmoid activation and residual summation. Our MGRConv is a balance between partial convolution and learnable attention maps, and boosts denoising performance better than other inpainting convolutions with similar or even less parameters, memory, and training/inference time. Extensive experiments show that our proposed plug-and-play MGRConv can assist blind-spot based denoising networks to reach promising results on both existing single-image based and dataset based benchmarks. Yuhongze Zhou, Liguang Zhou, Issam H. Laradji, Tin Lun Lam, Yangsheng Xu |
ICASSP | 3 |
| 2023 | Constraining Representations Yields Models That Know What They Don't Know
João Monteiro 0002, Pau Rodríguez, Pierre-André Noël, Issam H. Laradji, David Vázquez 0001 |
ICLR | 4 |
| 2023 | OCR-VQGAN: Taming Text-within-Image GenerationabstractSynthetic image generation has recently experienced significant improvements in domains such as natural image or art generation. However, the problem of figure and diagram generation remains unexplored. A challenging aspect of generating figures and diagrams is effectively rendering readable texts within the images. To alleviate this problem, we present OCR-VQGAN, an image encoder, and decoder that leverages OCR pre-trained features to optimize a text perceptual loss, encouraging the architecture to preserve high-fidelity text and diagram structure. To explore our approach, we introduce the Paper2Fig100k dataset, with over 100k images of figures and texts from research papers. The figures show architecture diagrams and methodologies of articles available at arXiv.org from fields like artificial intelligence and computer vision. Figures usually include text and discrete objects, e.g., boxes in a diagram, with lines and arrows that connect them. We demonstrate the effectiveness of OCR-VQGAN by conducting several experiments on the task of figure reconstruction. Additionally, we explore the qualitative and quantitative impact of weighting different perceptual metrics in the overall loss function. We release code, models, and dataset at https://github.com/joanrod/ocr-vqgan. Juan A. Rodríguez, David Vázquez 0001, Issam H. Laradji, Marco Pedersoli, Pau Rodríguez |
WACV | 3 |
| 2023 | A Survey of Self-Supervised and Few-Shot Object DetectionabstractLabeling data is often expensive and time-consuming, especially for tasks such as object detection and instance segmentation, which require dense labeling of the image. While few-shot object detection is about training a model on novel (unseen) object classes with little data, it still requires prior training on many labeled examples of base (seen) classes. On the other hand, self-supervised methods aim at learning representations from unlabeled data which transfer well to downstream tasks such as object detection. Combining few-shot and self-supervised object detection is a promising research direction. In this survey, we review and characterize the most recent approaches on few-shot and self-supervised object detection. Then, we give our main takeaways and discuss future research directions. Project page: https://gabrielhuang.github.io/fsod-survey/. Gabriel Huang, Issam H. Laradji, David Vázquez 0001, Simon Lacoste-Julien, Pau Rodríguez |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Consistency-CAM: Towards Improved Weakly Supervised Semantic Segmentation
Sai Rajeswar, Issam H. Laradji, Pau Rodríguez, David Vázquez 0001, Aaron C. Courville |
BMVC | 2 |
| 2022 | OSM: An Open Set Matting Framework with OOD Detection and Few-Shot Learning
Yuhongze Zhou, Issam H. Laradji, Liguang Zhou, Derek Nowrouzezahrai |
BMVC | 2 |
| 2022 | Kubric: A scalable dataset generatorabstractData is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and training details. But collecting, processing and annotating real data at scale is difficult, expensive, and frequently raises additional privacy, fairness and legal concerns. Synthetic data is a powerful tool with the potential to address these shortcomings: 1) it is cheap 2) supports rich ground-truth annotations 3) offers full control over data and 4) can circumvent or mitigate problems regarding bias, privacy and licensing. Unfortunately, software tools for effective data generation are less mature than those for architecture design and training, which leads to fragmented generation efforts. To address these problems we introduce Kubric, an open-source Python framework that interfaces with PyBullet and Blender to generate photo-realistic scenes, with rich annotations, and seamlessly scales to large jobs distributed over thousands of machines, and generating TBs of data. We demonstrate the effectiveness of Kubric by presenting a series of 13 different generated datasets for tasks ranging from studying 3D NeRF models to optical flow estimation. We release Kubric, the used assets, all of the generation code, as well as the rendered datasets for reuse and modification. Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J. Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, Thomas Kipf, Abhijit Kundu, Dmitry Lagun, Issam H. Laradji, Hsueh-Ti Derek Liu, Henning Meyer, Yishu Miao, Derek Nowrouzezahrai, A. Cengiz Öztireli, Etienne Pot, Noha Radwan, Daniel Rebain, Sara Sabour, Mehdi S. M. Sajjadi, Matan Sela, Vincent Sitzmann, Austin Stone, Deqing Sun, Suhani Vora, Tianhao Wu 0003, Kwang Moo Yi, Fangcheng Zhong, Andrea Tagliasacchi |
CVPR | 14 |
| 2022 | Neural Point Light FieldsabstractWe introduce Neural Point Light Fields that represent scenes implicitly with a light field living on a sparse point cloud. Combining differentiable volume rendering with learned implicit density representations has made it possible to synthesize photo-realistic images for novel views of small scenes. As neural volumetric rendering methods require dense sampling of the underlying functional scene representation, at hundreds of samples along a ray cast through the volume, they are fundamentally limited to small scenes with the same objects projected to hundreds of training views. Promoting sparse point clouds to neural implicit light fields allows us to represent large scenes effectively with only a single radiance evaluation per ray. These point light fields are as a function of the ray direction, and local point feature neighborhood, allowing us to interpolate the light field conditioned training images without dense object coverage and parallax. We assess the proposed method for novel view synthesis on large driving scenarios, where we synthesize realistic unseen views that existing implicit approaches fail to represent. We validate that Neural Point Light Fields make it possible to predict videos along unseen trajectories previously only feasible to generate by explicitly modeling the scene. Julian Ost, Issam H. Laradji, Alejandro Newell, Yuval Bahat, Felix Heide |
CVPR | 2 |
| 2022 | CVPR 2020 continual learning in computer vision competition: Approaches, results, current challenges and future directions
Vincenzo Lomonaco, Lorenzo Pellegrini, Pau Rodríguez, Massimo Caccia, Qi She, Quentin Jodelet, Ruiping Wang 0001, Zheda Mai, David Vázquez 0001, German Ignacio Parisi, Nikhil Churamani, Marc Pickett, Issam H. Laradji, Davide Maltoni |
Artif. Intell. | 14 |
| 2022 | Let's Make Block Coordinate Descent Converge Faster: Faster Greedy Rules, Message-Passing, Active-Set Complexity, and Superlinear ConvergenceabstractBlock coordinate descent (BCD) methods are widely used for large-scale numerical optimization because of their cheap iteration costs, low memory requirements, amenability to parallelization, and ability to exploit problem structure. Three main algorithmic choices influence the performance of BCD methods: the block partitioning strategy, the block selection rule, and the block update rule. In this paper we explore all three of these building blocks and propose variations for each that can significantly improve the progress made by each BCD iteration. We (i) propose new greedy block-selection strategies that guarantee more progress per iteration than the Gauss-Southwell rule; (ii) explore practical issues like how to implement the new rules when using "variable" blocks; (iii) explore the use of message-passing to compute matrix or Newton updates efficiently on huge blocks for problems with sparse dependencies between variables; and (iv) consider optimal active manifold identification, which leads to bounds on the "active-set complexity" of BCD methods and leads to superlinear convergence for certain problems with sparse solutions (and in some cases finite termination at an optimal solution). We support all of our findings with numerical results for the classic machine learning problems of least squares, logistic regression, multi-class logistic regression, label propagation, and L1-regularization. Julie Nutini, Issam H. Laradji, Mark Schmidt 0001 |
J. Mach. Learn. Res. | 2 |
| 2021 | Stochastic Polyak Step-size for SGD: An Adaptive Learning Rate for Fast ConvergenceabstractWe propose a stochastic variant of the classical Polyak step-size (Polyak, 1987) commonly used in the subgradient method. Although computing the Polyak step-size requires knowledge of the optimal function values, this information is readily available for typical modern machine learning applications. Consequently, the proposed stochastic Polyak step-size (SPS) is an attractive choice for setting the learning rate for stochastic gradient descent (SGD). We provide theoretical convergence guarantees for SGD equipped with SPS in different settings, including strongly convex, convex and non-convex functions. Furthermore, our analysis results in novel convergence guarantees for SGD with a constant step-size. We show that SPS is particularly effective when training over-parameterized models capable of interpolating the training data. In this setting, we prove that SPS enables SGD to converge to the true solution at a fast rate without requiring the knowledge of any problem-dependent constants or additional computational overhead. We experimentally validate our theoretical results via extensive experiments on synthetic and real datasets. We demonstrate the strong performance of SGD with SPS compared to state-of-the-art optimization methods when training over-parameterized models. Nicolas Loizou, Sharan Vaswani, Issam H. Laradji, Simon Lacoste-Julien |
AISTATS | 3 |
| 2021 | Beyond Trivial Counterfactual Explanations with Diverse Valuable ExplanationsabstractExplainability for machine learning models has gained considerable attention within the research community given the importance of deploying more reliable machine-learning systems. In computer vision applications, generative counterfactual methods indicate how to perturb a model’s input to change its prediction, providing details about the model’s decision-making. Current methods tend to generate trivial counterfactuals about a model’s decisions, as they often suggest to exaggerate or remove the presence of the attribute being classified. For the machine learning practitioner, these types of counterfactuals offer little value, since they provide no new information about undesired model or data biases. In this work, we identify the problem of trivial counterfactual generation and we propose DiVE to alleviate it. DiVE learns a perturbation in a disentangled latent space that is constrained using a diversity-enforcing loss to uncover multiple valuable explanations about the model’s prediction. Further, we introduce a mechanism to prevent the model from producing trivial explanations. Experiments on CelebA and Synbols demonstrate that our model improves the success rate of producing high-quality valuable explanations when compared to previous state-of-the-art methods. Code is available at https://github.com/ElementAI/beyond-trivial-explanations. Pau Rodríguez, Massimo Caccia, Alexandre Lacoste, Lee Zamparo, Issam H. Laradji, Laurent Charlin, David Vázquez 0001 |
ICCV | 5 |
| 2021 | A Weakly Supervised Consistency-based Learning Method for COVID-19 Segmentation in CT ImagesabstractCoronavirus Disease 2019 (COVID-19) has spread aggressively across the world causing an existential health crisis. Thus, having a system that automatically detects COVID-19 in tomography (CT) images can assist in quantifying the severity of the illness. Unfortunately, labelling chest CT scans requires significant domain expertise, time, and effort. We address these labelling challenges by only requiring point annotations, a single pixel for each infected region on a CT image. This labeling scheme allows annotators to label a pixel in a likely infected region, only taking 1-3 seconds, as opposed to 10-15 seconds to segment a region. Conventionally, segmentation models train on point-level annotations using the cross-entropy loss function on these labels. However, these models often suffer from low precision. Thus, we propose a consistency-based (CB) loss function that encourages the output predictions to be consistent with spatial transformations of the input images. The experiments on 3 open-source COVID-19 datasets show that this loss function yields significant improvement over conventional point-level loss functions and almost matches the performance of models trained with full supervision with much less human effort. Code is available at: https://github.com/IssamLaradji/covid19_weak_supervision. Issam H. Laradji, Pau Rodríguez, Oscar Mañas, Keegan Lensink, Marco Law, Lironne Kurzman, William Parker, David Vázquez 0001, Derek Nowrouzezahrai |
WACV | 1 |
| 2021 | Learning Data Augmentation with Online Bilevel Optimization for Image ClassificationabstractData augmentation is a key practice in machine learning for improving generalization performance. However, finding the best data augmentation hyperparameters requires domain knowledge or a computationally demanding search. We address this issue by proposing an efficient approach to automatically train a network that learns an effective distribution of transformations to improve its generalization. Using bilevel optimization, we directly optimize the data augmentation parameters using a validation set. This framework can be used as a general solution to learn the optimal data augmentation jointly with an end task model like a classifier. Results show that our joint training method produces an image classification accuracy that is comparable to or better than carefully hand-crafted data augmentation. Yet, it does not need an expensive external validation loop on the data augmentation hyperparameters. Saypraseuth Mounsaveng, Issam H. Laradji, Ismail Ben Ayed, David Vázquez 0001, Marco Pedersoli |
WACV | 2 |
| 2021 | A Deep Learning Localization Method for Measuring Abdominal Muscle Dimensions in Ultrasound ImagesabstractHealth professionals extensively use Two-Dimensional (2D) Ultrasound (US) videos and images to visualize and measure internal organs for various purposes including evaluation of muscle architectural changes. US images can be used to measure abdominal muscles dimensions for the diagnosis and creation of customized treatment plans for patients with Low Back Pain (LBP), however, they are difficult to interpret. Due to high variability, skilled professionals with specialized training are required to take measurements to avoid low intra-observer reliability. This variability stems from the challenging nature of accurately finding the correct spatial location of measurement endpoints in abdominal US images. In this paper, we use a Deep Learning (DL) approach to automate the measurement of the abdominal muscle thickness in 2D US images. By treating the problem as a localization task, we develop a modified Fully Convolutional Network (FCN) architecture to generate blobs of coordinate locations of measurement endpoints, similar to what a human operator does. We demonstrate that using the TrA400 US image dataset, our network achieves a Mean Absolute Error (MAE) of 0.3125 on the test set, which almost matches the performance of skilled ultrasound technicians. Our approach can facilitate next steps for automating the process of measurements in 2D US images, while reducing inter-observer as well as intra-observer variability for more effective clinical outcomes. Alzayat Saleh, Issam H. Laradji, Corey Lammie, David Vázquez 0001, Carol A. Flavell, Mostafa Rahimi Azghadi |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Fast and Furious Convergence: Stochastic Second Order Methods under InterpolationabstractWe consider stochastic second-order methods for minimizing smooth and strongly-convex functions under an interpolation condition satisfied by over-parameterized models. Under this condition, we show that the regularized subsampled Newton method (R-SSN) achieves global linear convergence with an adaptive step-size and a constant batch-size. By growing the batch size for both the subsampled gradient and Hessian, we show that R-SSN can converge at a quadratic rate in a local neighbourhood of the solution. We also show that R-SSN attains local linear convergence for the family of self-concordant functions. Furthermore, we analyze stochastic BFGS algorithms in the interpolation setting and prove their global linear convergence. We empirically evaluate stochastic L-BFGS and a "Hessian-free" implementation of R-SSN for binary classification on synthetic, linearly-separable datasets and real datasets under a kernel mapping. Our experimental results demonstrate the fast convergence of these methods, both in terms of the number of iterations and wall-clock time. Si Yi Meng, Sharan Vaswani, Issam H. Laradji, Mark Schmidt 0001, Simon Lacoste-Julien |
AISTATS | 3 |
| 2020 | Embedding Propagation: Smoother Manifold for Few-Shot Classification
Pau Rodríguez, Issam H. Laradji, Alexandre Drouin, Alexandre Lacoste |
ECCV (26) | 2 |
| 2020 | Looc: Localize Overlapping Objects with Count SupervisionabstractAcquiring count annotations generally requires less human effort than point-level and bounding box annotations. Thus, we propose the novel problem setup of localizing objects in dense scenes under this weaker supervision. We propose LOOC, a method to Localize Overlapping Objects with Count supervision. We train LOOC by alternating between two stages. In the first stage, LOOC learns to generate pseudo point-level annotations in a semi-supervised manner. In the second stage, LOOC uses a fully-supervised localization method that trains on these pseudo labels. The localization method is used to progressively improve the quality of the pseudo labels. We conducted experiments on popular counting datasets. For localization, LOOC achieves a strong new baseline in the novel problem setup where only count supervision is available. For counting, LOOC outperforms current state-of-the-art methods that only use count as their supervision. Issam H. Laradji, Rafael Pardinas, Pau Rodríguez, David Vázquez 0001 |
ICIP | 1 |
| 2020 | Proposal-Based Instance Segmentation With Point SupervisionabstractInstance segmentation methods often require costly per-pixel labels. We propose a method called WISE-Net that only requires point-level annotations. During training, the model only has access to a single pixel label per object, yet the task is to output full segmentation masks. To address this challenge, we construct a network with two branches: (1) a 10-calization network (L-Net) that predicts the location of each object; and (2) an embedding network (E-Net) that learns an embedding space where pixels of the same object are close. The segmentation masks for the located objects are obtained by grouping pixels with similar embeddings. We evaluate our approach on PASCAL VOC, COCO, KITTI and CityScapes datasets. The experiments show that our method (1) obtains competitive results compared to fully-supervised methods in certain scenarios; (2) outperforms fully-and weakly-supervised methods with a fixed annotation budget; and (3) establishes a first strong baseline for instance segmentation with point-level supervision. Issam H. Laradji, Negar Rostamzadeh, Pedro O. Pinheiro, David Vázquez 0001, Mark Schmidt 0001 |
ICIP | 1 |
| 2020 | Online Fast Adaptation and Knowledge Accumulation (OSAKA): a New Approach to Continual LearningabstractContinual learning agents experience a stream of (related) tasks. The main challenge is that the agent must not forget previous tasks and also adapt to novel tasks in the stream. We are interested in the intersection of two recent continual-learning scenarios. In meta-continual learning, the model is pre-trained using meta-learning to minimize catastrophic forgetting of previous tasks. In continual-meta learning, the aim is to train agents for faster remembering of previous tasks through adaptation. In their original formulations, both methods have limitations. We stand on their shoulders to propose a more general scenario, OSAKA, where an agent must quickly solve new (out-of-distribution) tasks, while also requiring fast remembering. We show that current continual learning, meta-learning, meta-continual learning, and continual-meta learning techniques fail in this new scenario. We propose Continual-MAML, an online extension of the popular MAML algorithm as a strong baseline for this scenario. We show in an empirical study that Continual-MAML is better suited to the new scenario than the aforementioned methodologies including standard continual learning and meta-learning approaches. Massimo Caccia, Pau Rodríguez, Oleksiy Ostapenko, Fabrice Normandin, Lucas Caccia, Issam H. Laradji, Irina Rish, Alexandre Lacoste, David Vázquez 0001, Laurent Charlin |
NeurIPS | 7 |
| 2020 | Synbols: Probing Learning Algorithms with Synthetic DatasetsabstractProgress in the field of machine learning has been fueled by the introduction of benchmark datasets pushing the limits of existing algorithms. Enabling the design of datasets to test specific properties and failure modes of learning algorithms is thus a problem of high interest, as it has a direct impact on innovation in the field. In this sense, we introduce Synbols — Synthetic Symbols — a tool for rapidly generating new datasets with a rich composition of latent features rendered in low resolution images. Synbols leverages the large amount of symbols available in the Unicode standard and the wide range of artistic font provided by the open font community. Our tool's high-level interface provides a language for rapidly generating new distributions on the latent features, including various types of textures and occlusions. To showcase the versatility of Synbols, we use it to dissect the limitations and flaws in standard learning algorithms in various learning setups including supervised learning, active learning, out of distribution generalization, unsupervised representation learning, and object counting. Alexandre Lacoste, Pau Rodríguez, Frederic Branchaud-Charron, Parmida Atighehchian, Massimo Caccia, Issam H. Laradji, Alexandre Drouin, Matt Craddock, Laurent Charlin, David Vázquez 0001 |
NeurIPS | 6 |
| 2019 | Where are the Masks: Instance Segmentation with Image-level Supervision
Issam H. Laradji, David Vázquez 0001, Mark Schmidt 0001 |
BMVC | 1 |
| 2019 | Efficient Deep Gaussian Process Models for Variable-Sized InputsabstractDeep Gaussian processes (DGP) have appealing Bayesian properties, can handle variable-sized data, and learn deep features. Their limitation is that they do not scale well with the size of the data. Existing approaches address this using a deep random feature (DRF) expansion model, which makes inference tractable by approximating DGPs. However, DRF is not suitable for variable-sized input data such as trees, graphs, and sequences. We introduce the GP-DRF, a novel Bayesian model with an input layer of GPs, followed by DRF layers. The key advantage is that the combination of GP and DRF leads to a tractable model that can both handle a variable-sized input as well as learn deep long-range dependency structures of the data. We provide a novel efficient method to simultaneously infer the posterior of GP's latent vectors and infer the posterior of DRF's internal weights and random frequencies. Our experiments show that GP-DRF outperforms the standard GP model and DRF model across many datasets. Furthermore, they demonstrate that GP-DRF enables improved uncertainty quantification compared to GP and DRF alone, with respect to a Bhattacharyya distance assessment. Issam H. Laradji, Mark Schmidt 0001, Vladimir Pavlovic 0001, Minyoung Kim 0001 |
IJCNN | 1 |
| 2019 | Painless Stochastic Gradient: Interpolation, Line-Search, and Convergence RatesabstractRecent works have shown that stochastic gradient descent (SGD) achieves the fast convergence rates of full-batch gradient descent for over-parameterized models satisfying certain interpolation conditions. However, the step-size used in these works depends on unknown quantities and SGD's practical performance heavily relies on the choice of this step-size. We propose to use line-search techniques to automatically set the step-size when training models that can interpolate the data. In the interpolation setting, we prove that SGD with a stochastic variant of the classic Armijo line-search attains the deterministic convergence rates for both convex and strongly-convex functions. Under additional assumptions, SGD with Armijo line-search is shown to achieve fast convergence for non-convex functions. Furthermore, we show that stochastic extra-gradient with a Lipschitz line-search attains linear convergence for an important class of non-convex functions and saddle-point problems satisfying interpolation. To improve the proposed methods' practical performance, we give heuristics to use larger step-sizes and acceleration. We compare the proposed algorithms against numerous optimization methods on standard classification tasks using both kernel methods and deep networks. The proposed methods result in competitive performance across all models and datasets, while being robust to the precise choices of hyper-parameters. For multi-class classification using deep networks, SGD with Armijo line-search results in both faster convergence and better generalization. Sharan Vaswani, Aaron Mishkin, Issam H. Laradji, Mark Schmidt 0001, Gauthier Gidel, Simon Lacoste-Julien |
NeurIPS | 3 |
| 2018 | Where Are the Blobs: Counting by Localization with Point Supervision
Issam H. Laradji, Negar Rostamzadeh, Pedro O. Pinheiro, David Vázquez 0001, Mark Schmidt 0001 |
ECCV (2) | 1 |
| 2018 | MASAGA: A Linearly-Convergent Stochastic First-Order Method for Optimization on Manifolds
Reza Babanezhad 0001, Issam H. Laradji, Alireza Shafaei, Mark Schmidt 0001 |
ECML/PKDD (2) | 2 |
| 2016 | Convergence Rates for Greedy Kaczmarz Algorithms, and Randomized Kaczmarz Rules Using the Orthogonality Graph
Julie Nutini, Behrooz Sepehry, Issam H. Laradji, Mark Schmidt 0001, Hoyt A. Koepke, Alim Virani |
UAI | 3 |
| 2015 | Coordinate Descent Converges Faster with the Gauss-Southwell Rule Than Random SelectionabstractThere has been significant recent work on the theory and application of randomized coordinate descent algorithms, beginning with the work of Nesterov [SIAM J. Optim., 22(2), 2012], who showed that a random-coordinate selection rule achieves the same convergence rate as the Gauss-Southwell selection rule. This result suggests that we should never use the Gauss-Southwell rule, as it is typically much more expensive than random selection. However, the empirical behaviours of these algorithms contradict this theoretical result: in applications where the computational costs of the selection rules are comparable, the Gauss-Southwell selection rule tends to perform substantially better than random coordinate selection. We give a simple analysis of the Gauss-Southwell rule showing that—except in extreme cases—it’s convergence rate is faster than choosing random coordinates. Further, in this work we (i) show that exact coordinate optimization improves the convergence rate for certain sparse problems, (ii) propose a Gauss-Southwell-Lipschitz rule that gives an even faster convergence rate given knowledge of the Lipschitz constants of the partial derivatives, (iii) analyze the effect of approximate Gauss-Southwell rules, and (iv) analyze proximal-gradient variants of the Gauss-Southwell rule. Julie Nutini, Mark Schmidt 0001, Issam H. Laradji, Michael P. Friedlander, Hoyt A. Koepke |
ICML | 3 |
| 2015 | Software defect prediction using ensemble learning on selected features
Issam H. Laradji, Mohammad R. Alshayeb, Lahouari Ghouti |
Inf. Softw. Technol. | 1 |
| 2014 | Sparse Single-Hidden Layer Feedforward Network for Mapping Natural Language Questions to SQL Queries
Issam H. Laradji, Lahouari Ghouti, Faisal Saleh, Musab AlTurki |
ICANN | 1 |
| 2013 | Perceptual hashing of color images using hypercomplex representationsabstractThis paper presents a new perceptual image hashing approach that exploits the image color information using hypercomplex (quaternionic) representations. Unlike grayscale-based techniques, the proposed approach preserves the color interaction between the image components that have a significant contribution in the generated perceptual image hash codes. Having a robust image hash function optimizes a wide range of applications including content-based retrieval, image authentication, and image watermarking. Initially, the input color image is processed in a “holistic” manner using the hypercomplex representation where the red, green and blue (RGB) components are handled as a single entity. Then, non-overlapping 8 × 8 image blocks are processed using the Quaternion Fourier transform (QFT). Binary image hash codes are generated by comparing the block mean frequency energy to the global mean frequency energy. For retrieval purposes, the Hamming distance (HD) is used as the comparison metric to retrieve perceptually similar images. The performance of the proposed perceptual hashing for color image is compared to that based on the conventional complex Fourier transform (CFT). Simulation results clearly indicate the superior retrieval performance of the proposed QFT-based perceptual hashing technique in term of HD values of intra-and inter-class image samples. Moreover, the performance improvement of the QFT-based technique is achieved at a computational complexity similar to the CFT-based scheme. Issam H. Laradji, Lahouari Ghouti, El-Hebri Khiari |
ICIP | 1 |