Projects
Image Super Resolution
Converting low resolution (128x128) to high resolution (512x512) images using Generative Adversarial Networks. Inspired by the inability to find a high resolution video for "The Diary of Jane" by Breaking Benjamin. The GAN learns to generate realistic high-res images by training a discriminator to distinguish real from generated images, while the generator learns to produce outputs indistinguishable from real high-res images.
Voice Style Transfer
Extracting the timbre of one voice and superimposing it on another, inspired by image style transfer. Uses two convolutional encoders — a deep encoder for content (phonemes) and a shallow encoder for style (timbre) — to minimize content loss against the original input and style loss against the target voice.
Audio Classification
Two classification projects: (1) Emotion classification from speech using the emo-db dataset, and (2) Guitar chords classification using a custom synthesized dataset. Both use MFCC feature extraction fed into deep convolutional networks to learn spectral and temporal audio features.
Image Enhancement with Distributed Deep Learning
Implemented a distributed encoder-decoder architecture to convert low-resolution images to high-resolution using Ray for distributed training. Achieved a PSNR of 23.1 and SSIM of 0.72.
Retail Clickstream Analysis and Prediction
Developed a distributed system using PySpark to analyze eCommerce behavior from 14GB of click stream data. Derived insights like category analysis, cart conversion ratio, and abandonment rate. Built and deployed a real-time ML classification model using Streamlit achieving 78.42% accuracy.
Topological Analysis of Prompts in CLIP Model
Developed a topological algorithm to smartly select prompts and improve zero-shot performance of CLIP by ~1.3% over standard prompt ensembling. Uses the Mapper Select Algorithm to cluster relevant prompts from a large sample space of natural language textual prompts for downstream tasks.
Quantum Convolutional Neural Network (QCNN)
Implemented a quantum convolutional neural network on the Pennylane framework to investigate the usefulness and practicality of quantum neural networks for speech classification. Trained a simple 2-layer quantum circuit on the Google Speech dataset, achieving 52.23% accuracy.