Skip to content

Satyam · AI/ML Engineer

I turn AI demos into systems people actually use.

I build RAG pipelines, multi-agent systems, and the evaluation harnesses that keep them honest — end to end, from embedding pipeline to Dockerized API. Three internships spent closing the gap between a notebook and a 99% uptime deployment.

Selected results · productionn = 3 internships
RAG query latencyCloudily Scripts
8.2s 1.7s−79%
Hallucination rateAsvix · DigiLab
18% 11%−39%
Deploy timeIPtechhub
2h 15m−87%
Uptime · 500+ daily queriesAsvix · DigiLab
99.2%SLA

Real numbers from shipped work — not vanity metrics. Details in experience.

Selected work

Things I've shipped.

Seven projects — RAG, agents, fine-tuning, ML, and the eval harnesses that keep them honest. Each one has a full case study.

Production RAG with a guardrail gate, an evaluator agent, and a reproducible RAGAS harness.

  • Python
  • FastAPI
  • FAISS
  • RAGAS
  • OpenTelemetry

LLM-as-judge that treats position, verbosity, and self-enhancement bias as first-class problems.

  • Python
  • FastAPI
  • SQLAlchemy
  • OpenAI
  • Anthropic

Explainable AI-text detection on Binoculars cross-perplexity — 100% AUROC on HC3, with a fairness audit.

  • Python
  • PyTorch
  • Transformers
  • Streamlit
  • Docker

End-to-end MLOps: a 3-model benchmark served via FastAPI with data-drift monitoring.

  • XGBoost
  • FastAPI
  • Docker
  • scikit-learn
  • Monitoring

Experience

Three internships, one throughline.

Cloud infra → production RAG → hybrid retrieval at scale. Each role pushed a real system closer to reliable.

Jan 2026 — Apr 2026

Remote

AI Developer Intern

Asvix

  • Built the embedding pipeline for DigiLab, an AI chatbot on a LangChain + FAISS + Neo4j hybrid RAG stack — 500+ daily queries at 99.2% uptime.
  • Cut hallucination rate from 18% to 11% with context-aware response modules and iterative prompt refinement.
  • Lifted medical-query relevance 23% using tuned FAISS retrieval plus BM25 hybrid search.
  • Wired in LLM evaluation metrics (faithfulness, context recall) to drive prompt iteration with data, not guesses.
  • LangChain
  • FAISS
  • Neo4j
  • Hybrid RAG
  • LLM Eval

Jun 2025 — Jul 2025

On-site

AI Chatbot Development Intern

Cloudily Scripts

  • Built a production RAG pipeline (FAISS IVF128, cross-encoder reranking, BM25 dense retrieval) over 100+ page PDFs at 91% accuracy — cut support tickets 35%.
  • Cut query latency 79% (8.2s → 1.7s) with parallel embedding and semantic chunking.
  • Dockerized the stack and shrank the image 60% (2.1 GB → 840 MB) for faster deploys.
  • Python
  • RAG
  • FAISS
  • Docker
  • BM25

May 2024 — Jul 2024

Remote

Cloud Engineering Intern

IPtechhub

  • Deployed containerized ML inference on AWS EC2 with auto-scaling — 500+ daily requests at 99.5% uptime, infra cost down 32%.
  • Automated CI/CD with GitHub Actions: deploy time down 87.5% (2h → 15m), cold start from 45s to 8s.
  • AWS
  • Docker
  • CI/CD
  • GitHub Actions

About

The short version.

I build AI systems that ship. Three internships taught me what the gap between a working demo and a real deployment actually looks like — retrieval that grounds answers, evals that catch regressions, and Docker images small enough to deploy.

At Asvix I built the RAG pipeline behind DigiLab, an educational chatbot handling 500+ daily queries. At Cloudily Scripts I cut query latency from 8.2s to 1.7s on a live PDF RAG system. At IPtechhub I automated deployments from two hours to fifteen minutes.

I also co-authored research on hybrid syntax detection — AST parsing plus a gradient-boosting classifier across five languages, now being prepared for IEEE submission.

My approach is boring on purpose: make it simple, make it work, then make it better.

Education
B.Tech, AI & ML — United Institute of Technology (2026)
Focus
RAG · multi-agent systems · LLM evaluation
Research
Hybrid syntax detection — IEEE submission in prep
Based
India · open to remote / global
Languages
Python, SQL · English, Hindi

Stack

What I work with.

Tools I've shipped to production — not a wishlist. Depth in the AI/LLM and backend rows.

Languages

  • Python
  • JavaScript
  • TypeScript
  • SQL

AI & LLM

  • RAG
  • Agentic AI
  • LangChain
  • LangGraph
  • CrewAI
  • Embeddings
  • RAGAS
  • LLM Evaluation
  • Prompt Engineering

Backend

  • FastAPI
  • REST APIs
  • Node.js
  • Microservices

Data & ML

  • scikit-learn
  • Pandas
  • NumPy
  • XGBoost
  • NLP

Vector & DB

  • FAISS
  • Pinecone
  • PostgreSQL
  • MongoDB
  • Neo4j

DevOps & Cloud

  • Docker
  • AWS
  • CI/CD
  • GitHub Actions
  • OpenTelemetry
  • Linux

Contact

Let's build something that survives production.

Open to full-time AI/ML roles and freelance work — remote or global. If you're hiring for RAG, agents, or eval-heavy systems, I'd like to hear about it.

Currently open to opportunities — usually reply within a day.