William J. Beksi

dblp:151/9378 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0001-5377-2627ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 7 first-author · 12 since 2021Systems, architecture and hardware · 15 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Energy-Based Open-Set Active Learning for Object Classification
Zongyao Lyu, William J. Beksi
ICPR (6)2
2026 Person Re-identification via Generalized Class Prototypes
Md Ahmed Al Muzaddid, William J. Beksi
ICPR (6)2
2025 Instance Segmentation-Based Hazard Detection with Lunar South Pole Lighting
abstract
This paper addresses rock hazard detection for in-situ resource utilization (ISRU) robotic navigation in the challenging visual environment of the lunar south pole (LSP). We evaluate three state-of-the-art instance segmentation mod-els-Mask R-CNN, YOLOv8, and SAM-using a novel, synthetically generated dataset that simulates LSP-specific illumination challenges at sun angles of$2.5^{\circ}, 5^{\circ}$, and 7.5°. Additionally, we evaluate these approaches in both up and downsun driving with low solar angle light. This study highlights the potential of deep learning-based approaches for improving ISRU operations by reliably identifying visual surface hazards, such as rocks, which may impede robotic navigation and excavation in future lunar missions.
Joseph M. Cloud, Bradley C. Buckles, Thomas J. Muller, William J. Beksi, Jason M. Schuler
ICRA4
2025 Vision-Based Movement Primitives for Lunar Hazard Avoidance
abstract
To support sustainable infrastructure on the Moon, NASA is developing the In-Situ Resource Utilization (ISRU) Pilot Excavator (IPEx) to extract and transport lunar regolith for processing and construction. During its mission, IPEx will execute various driving patterns, primarily cycling between excavation and unloading sites, with additional ma-neuvers such as circular traverses around the lander and raster scans for environmental mapping. In this work, dynamic move-ment primitives (DMPs) are used to represent these patterns. We augment the DMPs with a vision-based real-time obstacle avoidance system to navigate surface hazards, such as rocks, encountered during traversal. Our approach is evaluated in a high-fidelity simulation replicating the challenging environment of the lunar south pole to demonstrate IPEx's ability to adapt to surface hazards while fulfilling its operational tasks.
Joseph M. Cloud, William J. Beksi, Jason M. Schuler
ICRA2
2025 Secrets of Edge-Informed Contrast Maximization for Event-Based Vision
abstract
Event cameras capture the motion of intensity gradients (edges) in the image plane in the form of rapid asynchronous events. When accumulated in 2D histograms, these events depict overlays of the edges in motion, consequently obscuring the spatial structure of the generating edges. Contrast maximization (CM) is an optimization framework that can reverse this effect and produce sharp spatial structures that resemble the moving intensity gradients by estimating the motion trajectories of the events. Nonetheless, CM is still an underexplored area of research with avenues for improvement. In this paper, we propose a novel hybrid approach that extends CM from uni-modal (events only) to bi-modal (events and edges). We leverage the underpinning concept that, given a reference time, optimally warped events produce sharp gradients consistent with the moving edge at that time. Specifically, we formalize a correlation-based objective to aid CM and provide key insights into the incorporation of multiscale and multireference techniques. Moreover, our edge-informed CM method yields superior sharpness scores and establishes new state-of-the-art event optical flow benchmarks on the MVSEC, DSEC, and ECD datasets.
Pritam Karmokar, Quan H. Nguyen, William J. Beksi
WACV3
2025 Agent deception via polynomial path planning
abstract
Deceptive behavior involves an intelligent agent creating plans that conceal its true intentions while appearing to pursue different goals. It is crucial for inducing confusion in various applications, including security, military, and competitive environments, where the ability to conceal true intentions can lead to significant strategic advantages. Developing better artificial intelligence (AI) for adversarial environments or strategic decision-making scenarios (e.g., more realistic testing of human or AI decision-making capabilities in games and simulations) requires effective deceptive planning. In this paper, we propose a novel polynomial path planner that enables an agent to deceive its observers. Our contributions include the following: (i) we develop a framework for obtaining ambiguous functions; (ii) we introduce new deception metrics; (iii) we present a method for standardizing trajectories to enable shape-based comparisons independent of speed; (iv) we conduct a human survey to evaluate deception and its relation to goal recognition; and (v) we outperform the state of the art on a multitude of deception metrics. Furthermore, our findings show that using more complex functions and increasing the level of misdirection greatly enhances agent deception effectiveness.
Nolan Gutierrez, Brian M. Sadler, William J. Beksi
Eng. Appl. Artif. Intell.3
2024 Semi-Supervised Variational Adversarial Active Learning via Learning to Rank and Agreement-Based Pseudo Labeling
Zongyao Lyu, William J. Beksi
ICPR (1)2
2024 Few-Shot Fruit Segmentation via Transfer Learning
abstract
Advancements in machine learning, computer vision, and robotics have paved the way for transformative solutions in various domains, particularly in agriculture. For example, accurate identification and segmentation of fruits from field images plays a crucial role in automating jobs such as harvesting, disease detection, and yield estimation. However, achieving robust and precise infield fruit segmentation remains a challenging task since large amounts of labeled data are required to handle variations in fruit size, shape, color, and occlusion. In this paper, we develop a few-shot semantic segmentation framework for infield fruits using transfer learning. Concretely, our work is aimed at addressing agricultural domains that lack publicly available labeled data. Motivated by similar success in urban scene parsing, we propose specialized pre-training using a public benchmark dataset for fruit transfer learning. By leveraging pre-trained neural networks, accurate semantic segmentation of fruit in the field is achieved with only a few labeled images. Furthermore, we show that models with pre-training learn to distinguish between fruit still on the trees and fruit that have fallen on the ground, and they can effectively transfer the knowledge to the target fruit dataset.
Jordan A. James, Heather K. Manching, Amanda M. Hulse-Kemp, William J. Beksi
ICRA4
2024 NTrack: A Multiple-Object Tracker and Dataset for Infield Cotton Boll Counting
abstract
In agriculture, automating the accurate tracking of fruits, vegetables, and fiber is a very tough problem. The issue becomes extremely challenging in dynamic field environments. Yet, this information is critical for making day-to-day agricultural decisions, assisting breeding programs, and much more. To tackle this dilemma, we introduceNTrack, a novel multiple object tracking framework based on the linear relationship between the locations ofneighboringtracks.NTrackcomputes dense optical flow and utilizes particle filtering to guide each tracker. Correspondences between detections and tracks are found through data association via direct observations and indirect cues, which are then combined to obtain an updated observation. Our modular multiple object tracking system is independent of the underlying detection method, thus allowing for the interchangeable use of any off-the-shelf object detector. We show the efficacy of our approach on the task of tracking and counting infield cotton bolls. Experimental results show that our system exceeds contemporary tracking and cotton boll-based counting methods by a large margin. Furthermore, we publicly release thefirstannotated cotton boll video dataset to the research community.Note to Practitioners—This work is motivated by the need to provide highly-accurate estimates of the total number of cotton bolls across an entire farm. We provide a multiple object tracking framework for automating the counting process using a dynamic motion model that can reidentify severely occluded objects. This information is immensely beneficial to agronomists and breeders. For example, it can allow them to accelerate the selection of genotypes and identify cotton cultivars that exhibit tolerance to adverse environmental conditions (e.g., drought, poor soil quality, etc.) via yield prediction. Since the performance of our tracker is tied to the accuracy of the object detector, the ability to swap detectors is important for enhancing the usability of the system. Our dataset structure was modeled after similar multiple object tracking datasets for the purpose of increasing adoption among practitioners. To implement our system, it is assumed that the image/video capture device is of high resolution and data is acquired under ideal lighting conditions. The framework can be run on any commercial-off-the-shelf platform (e.g., ground-based robots, unmanned aerial vehicles, etc.), with sufficient computing and memory resources, using a vision-based sensor.
Md Ahmed Al Muzaddid, William J. Beksi
IEEE Trans Autom. Sci. Eng.2
2023 LIST: Learning Implicitly from Spatial Transformers for Single-View 3D Reconstruction
abstract
Accurate reconstruction of both the geometric and topological details of a 3D object from a single 2D image embodies a fundamental challenge in computer vision. Existing explicit/implicit solutions to this problem struggle to recover self-occluded geometry and/or faithfully reconstruct topological shape structures. To resolve this dilemma, we introduce LIST, a novel neural architecture that leverages local and global image features to accurately reconstruct the geometric and topological structure of a 3D object from a single image. We utilize global 2D features to predict a coarse shape of the target object and then use it as a base for higher-resolution reconstruction. By leveraging both local 2D features from the image and 3D features from the coarse prediction, we can predict the signed distance between an arbitrary point and the target surface via an implicit predictor with great accuracy. Furthermore, our model does not require camera estimation or pixel alignment. It provides an uninfluenced reconstruction from the input-view direction. Through qualitative and quantitative analysis, we show the superiority of our model in reconstructing 3D objects from both synthetic and real-world images against the state of the art. Our source code is publicly available to the research community [13].
Mohammad Samiul Arshad, William J. Beksi
ICCV2
2023 Intuitive Robot Integration via Virtual Reality Workspaces
abstract
As robots become increasingly prominent in di-verse industrial settings, the desire for an accessible and reliable system has correspondingly increased. Yet, the task of meaningfully assessing the feasibility of introducing a new robotic component, or adding more robots into an existing infrastructure, remains a challenge. This is due to both the logistics of acquiring a robot and the need for expert knowledge in setting it up. In this paper, we address these concerns by developing a purely virtual simulation of a robotic system. Our proposed framework enables natural human-robot interaction through a visually immersive representation of the workspace. The main advantages of our approach are the following: (i) independence from a physical system, (ii) flexibility in defining the workspace and robotic tasks, and (iii) an intuitive interaction between the operator and the simulated environment. Not only does our system provide an enhanced understanding of 3D space to the operator, but it also encourages a hands-on way to perform robot programming. We evaluate the effectiveness of our method in applying novel automation assignments by training a robot in virtual reality and then executing the task on a real robot.
Minh Q. Tram, Joseph M. Cloud, William J. Beksi
ICRA3
2023 Lunar Excavator Mission Operations Using Dynamic Movement Primitives
abstract
To support sustainable infrastructure on the Moon, NASA must leverage robots to extract lunar resources for in-situ processing and construction. As part of this effort, NASA is launching the in-situ resource utilization (ISRU) Pilot Excavator later this decade to validate a robotic regolith excavator based on the Regolith Advanced Surface Systems Operations Robot (RASSOR). RASSOR is designed to extract and transport regolith to meet the needs of ISRU architectures. During its mission, Pilot Excavator will be tasked with driving in test patterns to demonstrate the operational concept. One possible test pattern is a circular trajectory around the lander while avoiding surface hazards such as lunar rocks. To this end, we utilize dynamic movement primitives to represent navigation sequences as primitive trajectories. We introduce a novel obstacle avoidance parameter, which is configured to avoid rocks throughout testing exercises. We demonstrate the effectiveness our method in a newly developed simulation tool called the Simulated Excavation Environment for Lunar Operations (SEELO) using models based on the NASA RASSOR 2.0 excavator. Our results show that the robot is able to safety and robustly navigate the lunar surface with densely populated rock obstacles while retaining the desired circle pattern behavior.
Joseph M. Cloud, Minh Q. Tram, William J. Beksi, Michael A. DuPuis
IROS3
2023 Single Image Super-Resolution via a Dual Interactive Implicit Neural Network
abstract
In this paper, we introduce a novel implicit neural network for the task of single image super-resolution at arbitrary scale factors. To do this, we represent an image as a decoding function that maps locations in the image along with their associated features to their reciprocal pixel attributes. Since the pixel locations are continuous in this representation, our method can refer to any location in an image of varying resolution. To retrieve an image of a particular resolution, we apply a decoding function to a grid of locations each of which refers to the center of a pixel in the output image. In contrast to other techniques, our dual interactive neural network decouples content and positional features. As a result, we obtain a fully implicit representation of the image that solves the super-resolution problem at (real-valued) elective scales using a single model. We demonstrate the efficacy and flexibility of our approach against the state of the art on publicly available benchmark datasets.
Quan H. Nguyen, William J. Beksi
WACV2
2023 IPVNet: Learning implicit point-voxel features for open-surface 3D reconstruction
Mohammad Samiul Arshad, William J. Beksi
J. Vis. Commun. Image Represent.2
2023 Rule-Based Safe Probabilistic Movement Primitive Control via Control Barrier Functions
abstract
In this paper, we develop a novel and safe control design approach that takes demonstrations provided by a human teacher to enable a robot to accomplish complex manipulation scenarios in dynamic environments. First, an overall task is divided into multiple simpler subtasks that are more appropriate for learning and control objectives. Then, by collecting human demonstrations, the subtasks that require robot movement are modeled by probabilistic movement primitives (ProMPs). We also study two strategies for modifying the ProMPs to avoid collisions with environmental obstacles. Finally, we introduce a rule-base control technique by utilizing a finite-state machine along with a unique means of control design for ProMPs. For the ProMP controller, we propose control barrier and Lyapunov functions to guide the system along a trajectory within the distribution defined by a ProMP while guaranteeing that the system state never leaves more than a desired distance from the distribution mean. This allows for better performance on nonlinear systems and offers solid stability and known bounds on the system state. A series of simulations and experimental studies demonstrate the efficacy of our approach and show that it can run in real time.Note to Practitioners—This paper is motivated by the need to create a teach-by-demonstration framework that captures the strengths of movement primitives and verifiable, safe control. We provide a framework that learns safe control laws from a probability distribution of robot trajectories through the use of advanced nonlinear control that incorporates safety constraints. Typically, such distributions are stochastic, making it difficult to offer any guarantees on safe operation. Our approach ensures that the distribution of allowed robot trajectories is within an envelope of safety and allows for robust operation of a robot. Furthermore, using our framework various probability distributions can be combined to represent complex scenarios in the environment. It will benefit practitioners by making it substantially easier to test and deploy accurate, efficient, and safe robots in complex real-world scenarios. The approach is currently limited to scenarios involving static obstacles, with dynamic obstacle avoidance an avenue of future effort.
Mohammad Reza Davoodi, Asif Iqbal 0009, Joseph M. Cloud, William J. Beksi, Nicholas R. Gans
IEEE Trans Autom. Sci. Eng.4
2022 Variable Rate Compression for Raw 3D Point Clouds
abstract
In this paper, we propose a novel variable rate deep compression architecture that operates on raw 3D point cloud data. The majority of learning-based point cloud compression methods work on a downsampled representation of the data. Moreover, many existing techniques require training multiple networks for different compression rates to generate consolidated point clouds of varying quality. In contrast, our network is capable of explicitly processing point clouds and generating a compressed description at a comprehensive range of bitrates. Furthermore, our approach ensures that there is no loss of information as a result of the voxelization process and the density of the point cloud does not affect the encoder/decoder performance. An extensive experimental evaluation shows that our model obtains state-of-the-art results, it is computationally efficient, and it can work directly with point cloud data thus avoiding an expensive voxelized representation.
Md Ahmed Al Muzaddid, William J. Beksi
ICRA2
2021 Learning the Next Best View for 3D Point Clouds via Topological Features
abstract
In this paper, we introduce a reinforcement learning approach utilizing a novel topology-based information gain metric for directing the next best view of a noisy 3D sensor. The metric combines the disjoint sections of an observed surface to focus on high-detail features such as holes and concave sections. Experimental results show that our approach can aid in establishing the placement of a robotic sensor to optimize the information provided by its streaming point cloud data. Furthermore, a labeled dataset of 3D objects, a CAD design for a custom robotic manipulator, and software for the transformation, union, and registration of point clouds has been publicly released to the research community.
Christopher Collander, William J. Beksi, Manfred Huber
ICRA2
2021 Thermal Image Super-Resolution Using Second-Order Channel Attention with Varying Receptive Fields
Nolan Gutierrez, William J. Beksi
ICVS2
2020 A Progressive Conditional Generative Adversarial Network for Generating Dense and Colored 3D Point Clouds
abstract
In this paper, we introduce a novel conditional generative adversarial network that creates dense 3D point clouds, with color, for assorted classes of objects in an unsupervised manner. To overcome the difficulty of capturing intricate details at high resolutions, we propose a point transformer that progressively grows the network through the use of graph convolutions. The network is composed of a leaf output layer and an initial set of branches. Every training iteration evolves a point vector into a point cloud of increasing resolution. After a fixed number of iterations, the number of branches is increased by replicating the last branch. Experimental results show that our network is capable of learning and mimicking a 3D data distribution, and produces colored point clouds with fine details at multiple resolutions.
Mohammad Samiul Arshad, William J. Beksi
3DV2
2019 A topology-based descriptor for 3D point cloud modeling: Theory and experiments
William J. Beksi, Nikolaos Papanikolopoulos
Image Vis. Comput.1
2018 Signature of Topologically Persistent Points for 3D Point Cloud Description
abstract
We present the Signature of Topologically Persistent Points (STPP), a global descriptor that encodes topological invariants of 3D point cloud data. These topological invariants include the zeroth and first homology groups and are computed using persistent homology, a method for finding the features of a topological space at different spatial resolutions. STPP is a competitive 3D point cloud descriptor when compared to the state of art and is resilient to noisy sensor data. We demonstrate experimentally on a publicly available RGB-D dataset that STPP can be used as a distinctive signature, thus allowing for 3D point cloud processing tasks such as object detection and classification.
William J. Beksi, Nikolaos Papanikolopoulos
ICRA1
2016 3D point cloud segmentation using topological persistence
abstract
In this paper, we present an approach to segment 3D point cloud data using ideas from persistent homology theory. The proposed algorithms first generate a simplicial complex representation of the point cloud dataset. Next, we compute the zeroth homology group of the complex which corresponds to the number of connected components. Finally, we extract the clusters of each connected component in the dataset. We show that this technique has several advantages over state of the art methods such as the ability to provide a stable segmentation of point cloud data under noisy or poor sampling conditions and its independence of a fixed distance metric.
William J. Beksi, Nikolaos Papanikolopoulos
ICRA1
2016 3D region segmentation using topological persistence
abstract
A `region' is an important concept in interpreting 3D point cloud data since regions may correspond to objects in a scene. To correctly interpret 3D point cloud data, we need to partition the dataset into regions that correspond to objects or parts of an object. In this paper, we present a region growing approach that combines global (topological) and local (color, surface normal) information to segment 3D point cloud data. Using ideas from persistent homology theory, our algorithm grows a simplicial complex representation of the point cloud dataset. At each step in the growth process we compute the zeroth homology group of the complex, which corresponds to the number of connected components, and use color and surface normal statistics to build regions. Lastly, we extract out the segmented regions of the dataset. We show that this method provides a stable segmentation of point cloud data in the presence of noise and poorly sampled data, thus providing advantages over contemporary region-based segmentation techniques.
William J. Beksi, Nikolaos Papanikolopoulos
IROS1
2016 Covariance based point cloud descriptors for object detection and recognition
Duc Fehr, William J. Beksi, Dimitris Zermas, Nikolaos Papanikolopoulos
Comput. Vis. Image Underst.2
2015 Object classification using dictionary learning and RGB-D covariance descriptors
abstract
In this paper, we introduce a dictionary learning framework using RGB-D covariance descriptors on point cloud data for performing object classification. Dictionary learning in combination with RGB-D covariance descriptors provides a compact and flexible description of point cloud data. Furthermore, the proposed framework is ideal for updating and sharing dictionaries among robots in a decentralized or cloud network. This work demonstrates the increased performance of 3D object classification utilizing covariance descriptors and dictionary learning over previous results with experiments performed on a publicly available RGB-D database.
William J. Beksi, Nikolaos Papanikolopoulos
ICRA1
2015 CORE: A Cloud-based Object Recognition Engine for robotics
abstract
An object recognition engine needs to extract discriminative features from data representing an object and accurately classify the object to be of practical use in robotics. Furthermore, the classification of the object must be rapidly performed in the presence of a voluminous stream of data. These conditions call for a distributed and scalable architecture that can utilize a cloud computing infrastructure for performing object recognition. This paper introduces a Cloud-based Object Recognition Engine (CORE) to address these needs. CORE is able to train on large-scale datasets, perform classification of 3D point cloud data, and efficiently transfer data in a robotic network.
William J. Beksi, John Spruth, Nikolaos Papanikolopoulos
IROS1
2014 Occlusion alleviation through motion using a mobile robot
abstract
Object segmentation and classification is an important and difficult task in robotic vision. The task is complicated even further when the different objects are partially or completely occluded. Allowing a robot to take measurements from varying points of view can help in alleviating or completely removing occlusions. A robot equipped with an RGB-D sensor has the capability of searching for new and better points of view to facilitate object recognition. In this work, a motion control algorithm is designed and implemented on a mobile robot to facilitate object classification in RGB-D data of clustered objects.
Duc Fehr, William J. Beksi, Dimitris Zermas, Nikolaos Papanikolopoulos
ICRA2
2014 RGB-D object classification using covariance descriptors
abstract
In this paper, we introduce a new covariance based feature descriptor to be used on “colored” point clouds gathered by a mobile robot equipped with an RGB-D camera. Although many recent descriptors provide adequate results, there is not yet a clear consensus on how to best tackle “colored” point clouds. We present the notion of a covariance on RGB-D data. Covariances have not only been proven to be successful in image processing, but in other domains as well. Their main advantage is that they provide a compact and flexible description of point clouds. Our work is a first step towards demonstrating the usability of the concept of covariances in conjunction with RGB-D data. Experiments performed on an RGB-D database and compared to previous results show the increased performance of our method.
Duc Fehr, William J. Beksi, Dimitris Zermas, Nikolaos Papanikolopoulos
ICRA2
2014 Point cloud culling for robot vision tasks under communication constraints
abstract
In this paper, we present two real-time methods for controlling data transmission in a robotic network that utilizes a remote computing infrastructure. The proposed algorithms use information and communication theory concepts to perform a highly efficient transfer of RGB-D data from a client (robot) to a server (cloud). We show that this approach makes it possible to conserve bandwidth and reduce network latency while allowing a mobile robot to perform vision tasks.
William J. Beksi, Nikolaos Papanikolopoulos
IROS1