Skip to main content
Open to DS / AI-ML roles

Data Scientist & AI/ML Engineer.

I turn messy, real-world data into models that make decisions. At BlackBox I cut document retrieval time 98% across 50,000+ documents; at PwC I hit 92% forecast accuracy on $20M/month in petrochemical trades. My focus is statistical rigor and pipelines built to be reproducible, not just to work once.

B.Sc. Computer Science, UMass Amherst ('26) · 3.82 GPA · Azure DP-100 & AWS AI Practitioner certified

Tech Stack

Tools I reach for across both sides of the stack.

Languages Frameworks / Libraries Data, Cloud & Analytics
Python SQL Java JavaScript TypeScript PyTorch scikit-learn XGBoost pandas NumPy statsmodels scipy Flask FastAPI React Node.js AWS SageMaker Lambda ECS Azure Docker Kubernetes PostgreSQL Snowflake Spark Tableau

Drag to rotate

How I Build Models

From raw data to a monitored, production endpoint — scroll to walk the pipeline.

STAGE 01

Ingest

Raw data from APIs, warehouses & streams

STAGE 02

Clean & Engineer

Handle nulls, build features

STAGE 03

Train & Validate

XGBoost, Random Forest, cross-validation

STAGE 04

Evaluate

Score against holdout & business metrics

STAGE 05

Deploy & Monitor

Serve predictions, track drift

Projects

Filter by discipline, or see everything at once.

EquiSight

Data Science

Challenge: walk-forward validated ARIMA vs. gradient-boosted (XGBoost) forecasting across 59 equities — 33.6% RMSE improvement at p ≈ 3×10⁻⁹ — then served it through a FastAPI + PostgreSQL backend and a React/Plotly dashboard, deployed live with sub-second API response times.

FastAPI PostgreSQL React / Plotly

NiftyCorridor

Data Science

Challenge: build a NIFTY50 options backtester where the overfitting controls are the product, not an afterthought — a parameter-sweep engine that ranks configs on a train window, re-validates only the top candidates out-of-sample, and computes Deflated Sharpe Ratio plus CSCV-based Probability of Backtest Overfitting on every shortlist before a config reaches the leaderboard. Self-hosted via Docker + nginx, gated behind Basic Auth rather than left public.

Python Streamlit Pydantic Docker

PolicyLens

AI/ML

Challenge: built the eval harness and a hand-verified 42-question golden set before writing a single retriever, then ran a 3-stage retrieval ablation (BM25 → dense → RRF hybrid fusion) over 109 SEC filings, NAIC model laws, and state insurance bulletins, scored with bootstrap confidence intervals. Hybrid fusion led on every metric (recall@10 0.767, MRR 0.495, nDCG@10 0.556); refusal accuracy held at 91.7%, and hand-tracing every false refusal showed all 9 were retrieval-recall gaps, not generation failures.

Python FastAPI Hybrid RAG Docker

COVID-19 Literature Clustering

AI/ML

Challenge: built an NLP pipeline (TF-IDF, t-SNE, K-means, LDA) to cluster 32,000+ CORD-19 research papers, preserving 95% variance while reducing a multi-thousand-paper corpus into navigable topic groups. (Lumiere Education Research)

Python TF-IDF t-SNE K-means / LDA

Document Retrieval & Fitment Scoring

AI/ML

Challenge: cut document retrieval time 98% (30 min to <1 min) across 50,000+ documents with a statistical ranking pipeline, and automated resume fitment scoring for ~500 applications/cycle with a TF-IDF-based NLP model — reducing screening time 60%.

Python NLP TF-IDF Statistical Scoring
Proprietary — built during internship at BlackBox

Petrochemical Price Forecasting

Data Science

Challenge: hit 92% forecast accuracy on $20M/month in petrochemical (Paraxylene) trades — cutting forecast error from 45% to 28% with XGBoost + Random Forest — then containerized the model with Docker and deployed on AWS SageMaker, scaling real-time inference to 10,000+ records/day.

XGBoost scikit-learn Docker AWS SageMaker
Proprietary — built during internship at PwC

Contact

Based in New York, NY — open to Data Science / AI-ML roles.