Partha Ghosh

dblp:66/8038 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2025 RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-Plane Networks
abstract
We present a compute and data-efficient video generative model designed to address long-term spatial and temporal dependencies. To capture long spatio-temporal dependencies, our approach incorporates a hybrid explicit-implicit tri-plane representation inspired by 3D-aware generative frameworks. Individual video frames are synthesized from an intermediate tri-plane representation, which is derived from one single latent code. This novel strategy more than halves the computational complexity measured in FLOPs compared to the most efficient state-of-the-art methods. Consequently, our approach facilitates the efficient and temporally coherent generation of videos. Moreover, our joint frame modeling approach, in contrast to autoregressive methods, mitigates the generation of visual artifacts. We further enhance the model’s capabilities by integrating an optical flow-based module within our Generative Adversarial Network (GAN) based generator architecture, thereby compensating for the constraints imposed by a smaller generator size. As a result, our model synthesizes high-fidelity video clips at a resolution of 256×256 pixels, with durations extending to more than 5 seconds at a frame rate of 30 fps. The efficacy and versatility of our approach are empirically validated through qualitative and quantitative assessments across three different datasets comprising both synthetic and real video clips. We will make our training and inference code public.
Partha Ghosh, Soubhik Sanyal, Cordelia Schmid, Bernhard Schölkopf
ICIP1
2025 Lagged Co-movement Prediction of Sectoral Indices in Stock Market using Frequent Itemset Mining
abstract
Stock price prediction has become a critical area of interest for investors and market analysts, though forecasting stock market trends remains a challenging endeavor due to the inherent volatility and unpredictability of the market. The process of stock price prediction typically involves estimating future prices based on historical data, market trends, and various socioeconomic factors. However, factors like market fluctuations, incomplete or erroneous data, and investor behavior add complexity to these predictions. Several methods are employed for stock price forecasting, including fundamental analysis, technical analysis, and machine learning approaches such as Linear Regression, Random Forest, and Long Short-Term Memory (LSTM) networks. This study focuses on using sectoral indices as benchmarking tools to evaluate sector performance. Specifically, it explores the co-movements of thirteen NSE sectoral indices, with one index chosen as the target. The analysis centers on using closing prices to measure sector performance and calculates the correlations between the target index and others. The six most highly correlated indices are identified, and association rule mining is used to uncover the relationships between these indices and the target index. The study aims to: (i) examine the interdependencies between the target sector and other sectors, and (ii) generate predictive rules for a sector’s performance based on the behavior of correlated sectors, providing valuable insights for making informed investment decisions.
Anjan Dutta 0002, Giridhar Maji, Partha Ghosh, Punyasha Chatterjee, Takaaki Goto, Soumya Sen 0001
SERA3
2024 SCULPT: Shape-Conditioned Unpaired Learning of Pose-dependent Clothed and Textured Human Meshes
abstract
We present SCULPT, a novel 3D generative model for clothed and textured 3D meshes of humans. Specifically, we devise a deep neural network that learns to represent the geometry and appearance distribution of clothed human bodies. Training such a model is challenging, as datasets of textured 3D meshes for humans are limited in size and accessibility. Our key observation is that there exist medium-sized 3D scan datasets like CAPE, as well as large-scale 2D image datasets of clothed humans and multiple appearances can be mapped to a single geometry. To effectively learn from the two data modalities, we propose an unpaired learning procedure for pose-dependent clothed and textured human meshes. Specifically, we learn a pose-dependent geometry space from 3D scan data. We represent this as per vertex displacements w.r.t. the SMPL model. Next, we train a geometry conditioned texture generator in an unsupervised way using the 2D image data. We use intermediate activations of the learned geometry model to condition our texture generator. To alleviate entanglement between pose and clothing type, and pose and clothing appearance, we condition both the texture and geometry generators with attribute labels such as clothing types for the geometry, and clothing colors for the texture generator. We automatically generated these conditioning labels for the 2D images based on the visual question answering model BLIP and CLIP. We validate our method on the SCULPT dataset, and compare to state-of-the-art 3D generative models for clothed human bodies. Our code and data can be found at https://sculpt.is.tue.mpg.de.
Soubhik Sanyal, Partha Ghosh, Michael J. Black, Justus Thies, Timo Bolkart
CVPR2
2024 A Machine Learning Based Automated Model for Managing Student Dropout
abstract
Addressing the persistent challenge of student dropout, particularly prevalent in developing countries like India, Bangladesh, etc. are of paramount importance. Factors such as poverty, natural calamities, and early marriages exacerbate this issue. High student dropout rates can negatively impact a country by diminishing its economic productivity, increasing social inequalities, and perpetuating a cycle of poverty. Addressing dropout issues requires comprehensive strategies to ensure a skilled and educated workforce, fostering societal well-being and global competitiveness. This research focuses on analysing comprehensive data on students who have dropped out. Thereafter, a machine learning based methodology is used to discern the underlying causes of student attrition in various schools. Furthermore, it allows for efficient monitoring of the state's educational landscape, with the ability to drill down to granular levels when necessary to identify specific regional challenges. The effectiveness of this approach is validated through the utilization of real-world datasets.
Partha Ghosh, Arnab Charit, Hindol Banerjee, Debanwesa Bandhu, Agniv Ghosh, Ankita Pal, Takaaki Goto, Soumya Sen 0001
SERA1
2024 Need of Public-Private Healthcare Collaboration for Managing Seasonal Dengue Fever in West Bengal
abstract
Dengue fever is mostly prevalent in tropical and subtropical regions, where Aedes mosquitoes, the primary vectors for the virus, thrive in warm and humid environments. In West Bengal, a province in India, the typical duration of the dengue disease spans two to three months, necessitating substantial infrastructure for dengue patients during this period. If the government heavily invests in developing this infrastructure, there's a risk of these facilities remaining underutilized during periods of low dengue incidence. Conversely, without adequate investment, dengue could potentially escalate into an epidemic. This research seeks to identify regions where insufficient infrastructure impedes public healthcare systems from catering to dengue patients. The primary focus of this paper is to evaluate the necessary degree of public-private collaborations needed to address seasonal dengue epidemics and pinpoint specific durations within the healthcare system that require attention.
Anwesha Nag, Takaaki Goto, Subhankar Roy, Partha Ghosh
SERA4
2024 Adversarial Likelihood Estimation With One-Way Flows
abstract
Generative Adversarial Networks (GANs) can produce high-quality samples, but do not provide an estimate of the probability density around the samples. However, it has been noted that maximizing the log-likelihood within an energy-based setting can lead to an adversarial framework where the discriminator provides unnormalized density (often called energy). We further develop this perspective, incorporate importance sampling, and show that 1) Wasserstein GAN performs a biased estimate of the partition function, and we propose instead to use an unbiased estimator; and 2) when optimizing for likelihood, one must maximize generator entropy. This is hypothesized to provide a better mode coverage. Different from previous works, we explicitly compute the density of the generated samples. This is the key enabler to designing an unbiased estimator of the partition function and computation of the generator entropy term. The generator density is obtained via a new type of flow network, called one-way flow network, that is less constrained in terms of architecture, as it does not require a tractable inverse function. Our experimental results show that our method converges faster, produces comparable sample quality to GANs with similar architecture, successfully avoids over-fitting to commonly used datasets and produces smooth low-dimensional latent representations of the training data.
Omri Ben-Dov, Pravir Singh Gupta, Victoria Fernández Abrevaya, Michael J. Black, Partha Ghosh
WACV5
2023 Scientific Organization of Blood Donation Camp Through Lexicographic Optimization and Taxicab Path Computation
abstract
Blood is the indispensable circulating fluid for sustaining human life. On demand supply of quality blood is a big challenge for every government in all developing countries. Specially, in festive seasons and winter, supplying quality blood on time is a big medical challenge. On the other hand, the consequences of mismanaged blood donation camp may lead to excess supply of human blood units. Also, in some cases, it is being noticed that human blood units are getting corrupted in transit from the blood donation camp to the blood bank. Hence, several units of human blood are getting spoiled over the time due to mismanagement and/or maintenance. In this research, we have applied a lexicographic optimization based model for finding best available blood bank from the point of blood donation camp. Alternative taxicab geometry based paths are used for finding best possible shortest path from the blood donation camp to the blood bank.
Partha Ghosh, Takaaki Goto, Leena Jana Ghosh, Soumya Sen 0001
SERA1
2023 Detection of a Novel Object-Detection-Based Cheat Tool for First-Person Shooter Games Using Machine Learning
abstract
Detection of novel game cheating tools is critical for ensuring fair online play. Such cheating tools are visual-based and effectively avoid detection because they do not change the data of game software. With the development and popularity of artificial intelligence technology, it has become easier for individuals to develop cheating tools, such as a new cheating tool for first-person shooter games that searches for characters on the game screen and automatically targets them. Therefore, in this study, a new cheat detection method is proposed using machine learning. The proposed method can be used to detect new cheating tools based on object detection.
Zhang Xiao, Takaaki Goto, Partha Ghosh, Tadaaki Kirishima, Kensei Tsuchida
SERA3
2021 Populating 3D Scenes by Learning Human-Scene Interaction
abstract
Humans live within a 3D space and constantly interact with it to perform tasks. Such interactions involve physical contact between surfaces that is semantically meaningful. Our goal is to learn how humans interact with scenes and leverage this to enable virtual characters to do the same. To that end, we introduce a novel Human-Scene Interaction (HSI) model that encodes proximal relationships, called POSA for "Pose with prOximitieS and contActs". The representation of interaction is body-centric, which enables it to generalize to new scenes. Specifically, POSA augments the SMPL-X parametric human body model such that, for every mesh vertex, it encodes (a) the contact probability with the scene surface and (b) the corresponding semantic scene label. We learn POSA with a VAE conditioned on the SMPL-X vertices, and train on the PROX dataset, which contains SMPL-X meshes of people interacting with 3D scenes, and the corresponding scene semantics from the PROX-E dataset. We demonstrate the value of POSA with two applications. First, we automatically place 3D scans of people in scenes. We use a SMPL-X model fit to the scan as a proxy and then find its most likely placement in 3D. POSA provides an effective representation to search for "affordances" in the scene that match the likely contact relationships for that pose. We perform a perceptual study that shows significant improvement over the state of the art on this task. Second, we show that POSA’s learned representation of body-scene interaction supports monocular human pose estimation that is consistent with a 3D scene, improving on the state of the art. Our model and code are available for research purposes at https://posa.is.tue.mpg.de.
Mohamed Hassan 0003, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas, Michael J. Black
CVPR2
2020 GIF: Generative Interpretable Faces
abstract
Photo-realistic visualization and animation of expressive human faces have been a long standing challenge. 3D face modeling methods provide parametric control but generates unrealistic images, on the other hand, generative 2D models like GANs (Generative Adversarial Networks) output photo-realistic face images, but lack explicit control. Recent methods gain partial control, either by attempting to disentangle different factors in an unsupervised manner, or by adding control post hoc to a pre-trained model. Unconditional GANs, however, may entangle factors that are hard to undo later. We condition our generative model on pre-defined control parameters to encourage disentanglement in the generation process. Specifically, we condition StyleGAN2 on FLAME, a generative 3D face model. While conditioning on FLAME parameters yields unsatisfactory results, we find that conditioning on rendered FLAME geometry and photometric details works well. This gives us a generative 2D face model named GIF (Generative Interpretable Faces) that offers FLAME's parametric control. Here, interpretable refers to the semantic meaning of different parameters. Given FLAME parameters for shape, pose, expressions, parameters for appearance, lighting, and an additional style vector, GIF outputs photo-realistic face images. We perform an AMT based perceptual study to quantitatively and qualitatively evaluate how well GIF follows its conditioning. The code, data, and trained model are publicly available for research purposes at http://gif.is.tue.mpg.de.
Partha Ghosh, Pravir Singh Gupta, Roy Uziel, Anurag Ranjan, Michael J. Black, Timo Bolkart
3DV1
2020 From Variational to Deterministic Autoencoders
Partha Ghosh, Mehdi S. M. Sajjadi, Antonio Vergari, Michael J. Black, Bernhard Schölkopf
ICLR1
2020 An Improved Intrusion Detection System to Preserve Security in Cloud Environment
abstract
Cloud computing, also known as on-demand computing, provides different kinds of services for the users. As the name suggests, its increasing demand makes it prone to various intruders affecting the privacy and integrity of the data stored in the cloud. To cope with this situation, intrusion detection systems (IDS) are implemented in the cloud. An effective IDS constitutes of less time-consuming algorithm with less space complexity and higher accuracy. To do so, the number of features are reduced while maintaining minimal loss of information. In this paper, the authors have proposed a model by which the features are selected on the basis of mutual information gain among correlated features. To achieve this, they first group the features according to the correlativity. Then from each group, the features with the highest mutual information gain in their respective groups are selected. This led them to a reduced feature set which provides quick learning and thus produces a better IDS that would secure the data in the cloud.
Partha Ghosh, Sumit Biswas, Shivam Shakti, Santanu Phadikar
Int. J. Inf. Secur. Priv.1
2019 Resisting Adversarial Attacks Using Gaussian Mixture Variational Autoencoders
abstract
Susceptibility of deep neural networks to adversarial attacks poses a major theoretical and practical challenge. All efforts to harden classifiers against such attacks have seen limited success till now. Two distinct categories of samples against which deep neural networks are vulnerable, “adversarial samples” and “fooling samples”, have been tackled separately so far due to the difficulty posed when considered together. In this work, we show how one can defend against them both under a unified framework. Our model has the form of a variational autoencoder with a Gaussian mixture prior on the latent variable, such that each mixture component corresponds to a single class. We show how selective classification can be performed using this model, thereby causing the adversarial objective to entail a conflict. The proposed method leads to the rejection of adversarial samples instead of misclassification, while maintaining high precision and recall on test data. It also inherently provides a way of learning a selective classifier in a semi-supervised scenario, which can similarly resist adversarial attacks. We further show how one can reclassify the detected adversarial samples by iterative optimization.1
Partha Ghosh, Arpan Losalka, Michael J. Black
AAAI1
2018 Chaotic firefly algorithm-based fuzzy C-means algorithm for segmentation of brain tissues in magnetic resonance images
Partha Ghosh, Kalyani Mali, Sitansu Kumar Das
J. Vis. Commun. Image Represent.1
2017 Learning Human Motion Models for Long-Term Predictions
abstract
We propose a new architecture for the learning of predictive spatio-temporal motion models from data alone. Our approach, dubbed the Dropout Autoencoder LSTM (DAELSTM), is capable of synthesizing natural looking motion sequences over long-time horizons1 without catastrophic drift or motion degradation. The model consists of two components, a 3-layer recurrent neural network to model temporal aspects and a novel autoencoder that is trained to implicitly recover the spatial structure of the human skeleton via randomly removing information about joints during training. This Dropout Autoencoder (DAE) is then used to filter each predicted pose by a 3-layer LSTM network, reducing accumulation of correlated error and hence drift over time. Furthermore to alleviate insufficiency of commonly used quality metric, we propose a new evaluation protocol using action classifiers to assess the quality of synthetic motion sequences. The proposed protocol can be used to assess quality of generated sequences of arbitrary length. Finally, we evaluate our proposed method on two of the largest motion-capture datasets available and show that our model outperforms the state-of-the-art techniques on a variety of actions, including cyclic and acyclic motion, and that it can produce natural looking sequences over longer time horizons than previous methods.
Partha Ghosh, Jie Song 0006, Emre Aksan, Otmar Hilliges
3DV1
2016 A novel technique for interdependent trim code optimization
abstract
As integrated circuits become increasingly complex, the fabrication process renders them more marginal to specification parameters. Consequently, these circuits have to be tuned post fabrication to compensate for process marginality and thereby optimize their performance as well as to improve yields. This process of tuning is commonly termed as trimming, wherein the right set of digital trim codes is identified and used as a calibration setting by writing these codes into the hardware configuration registers. This paper addresses the problem of carrying out multi-variable trims involving two or more parameters, codes for which must be simultaneously searched and set, in order to attain the desired performance of the circuit. A novel low-cost hardware implementation of the Simplex minimization algorithm with speed-up improvements is described. This implementation is amenable for on-chip BIST (built-in self-test) and can also be directly implemented as part of the ATE (automatic test equipment) program. Experimental results are presented on two industrial circuits. Improvements in terms of trim code search time and attaining better performance are shown.
Pankaj Bongale, Vinothkumar Sundaresan, Partha Ghosh, Rubin A. Parekhji
VTS3
2010 Fuzzy graph representation of a fuzzy concept lattice
Partha Ghosh, Krishna Kundu, Debasis Sarkar
Fuzzy Sets Syst.1