PRESENT
ML Intern
Building NLP and RAG systems. Currently developing a document analyzer and classifier on bart-large-mnli, using zero-shot inference to label document types without task-specific training data.
ML Intern @ Cubeage Technologies Pvt. Ltd. — building NLP and RAG systems, currently a zero-shot document analyzer and classifier on bart-large-mnli. Also developing OilTrace, a satellite-based spill origin tracer that detects oil spills in SAR imagery and backtraces drift to the source using wind, current, tide and sea-surface temperature. Previously ML intern @ Tantrata Digital.
Where I come from, what I work on, and what I'm optimising for.
I'm a computer engineering student and software intern based in Pune, working at the intersection of applied machine learning and pragmatic backend systems.
My work sits between two disciplines. On one side, training and evaluating models — currently transformer-based NLP: zero-shot document classification with NLI models, and the evaluation harnesses that tell you whether it actually works. Retrieval is the newer ground I am working through. On the other, the unglamorous plumbing that makes them useful: ETL, schema design in PostgreSQL, deployment on Vercel/Supabase.
I write C++ when the problem calls for it — graph algorithms, competitive programming, anything where the constant factor matters. Python is where I think; C++ is where I measure.
Ordered by comfort and time-under-the-hood, not by hype. Row 1 is where I ship daily.
Three internships to date, one active. Compact log — reverse chronological.
Building NLP and RAG systems. Currently developing a document analyzer and classifier on bart-large-mnli, using zero-shot inference to label document types without task-specific training data.
Contributed to specialised software cycles focused on data-driven logic inside AI/ML workflows. Built feature-extraction pipelines, model evaluation harnesses, and internal tooling in Python.
Worked on digital-transformation projects — data cleaning, business-process modelling, and prototype dashboards. Focused summer program.
A short set of projects that show how I think — across ML, systems, and web.
SIH project. Analyses satellite imagery to detect oil spills, then backtraces the drift to the actual source using particle physics and environmental factors: wind speed, sea-surface temperature, water flow, tide and current.
A document analyzer that reads any long-form document — land registration, rental agreements, resumes — and returns the important terms, clauses and risks as structured data. For resumes it produces an ATS-style score. Runs on local LLMs (Qwen 2.5 7B via Ollama) with the Google Gemini API as a fallback.
Tracks health via day-to-day signals like steps and vitals, surfaces insights over time, and analyses reported symptoms to predict likely conditions. React front-end with Supabase for auth and row-level-secure persistence.
Python + SQL system for donor registration, inventory tracking, and request routing across hospital records. Focus on schema correctness, transactional integrity, and a clean admin surface.
Current work at Cubeage Technologies. Labels long-form documents by type — contracts, notices, invoices, identity records — by turning each candidate label into a natural-language hypothesis and scoring it against the document with bart-large-mnli's entailment head. Because the classification is zero-shot, supporting a new document type means editing the label set rather than collecting examples and retraining.
Available for remote collaborations and internships. I read every message.