MK

Hi, I'm

MuzzamilKhalid

AI ENGINEER

AI Engineer specializing in agentic and generative AI. I build and ship production LLM systems: multi-agent orchestration (LangGraph, MCP), agentic RAG, real-time voice, inference optimization, and evals/observability. Several live products, including a monetized SaaS with 150+ daily active users.

0+ daily active users
across shipped products

Available for
opportunities

agent · run()live
user
run statsrunning…
steps 1tools 0lat 0.5s443 / 8k tok
ask the agent
LangGraph · MCP · tool-callingReAct
02

ABOUT ME

The engineer
behind the work.

I'm an AI Engineer based in Karachi, specializing in agentic LLM systems. Day to day that means LangGraph and MCP for multi-agent orchestration, agentic RAG over vector databases, LiveKit for real-time voice, and FastAPI services shipped with Docker and CI/CD pipelines.

I graduated in Artificial Intelligence from DUET in 2026, and interned at Logictech Digital building production RAG pipelines. Alongside that I've shipped four live products, including a monetized SaaS with 150+ daily active users and a GPU video pipeline where a quantized model cut inference cost by roughly half. The part I enjoy most is what happens after launch, when a system has to keep working.

0Live products shipped
0+Daily active users
0+Monthly subscribers
0Internships
03

SELECTED WORK

Selected
work.

Six systems I built and shipped: multi-agent orchestration, distributed GPU pipelines, MCP runtimes and production SaaS. Open any card for the architecture and the trade-offs behind it.

LIVE DEMO · 8×
MULTI-AGENT · REAL-TIME VOICEFINAL-YEAR PROJECT

CosMentor

My final-year project: a real-time, two-way voice tutor that teaches live coding across Python and JavaScript through 40+ structured lessons, pairing spoken conversation with a shared in-browser IDE. Three agents coordinate behind it over a full-duplex speech loop.

3 coordinated agents · 40+ lessons · Python & JavaScript

LiveKit · Qdrant · FastAPI · Neon Postgres · Pyodide WASM · React 19

LIVE DEMO · 3×
PRODUCTION SAAS

realistichandwriting.com

A browser-side rendering and OCR pipeline serving a paying subscriber base in production.

150+ daily active users · 80+ monthly subscribers

Next.js · Firebase · Stripe + Polar webhooks · Tesseract WASM · TipTap · jsPDF

LIVE DEMO · 5×
DISTRIBUTED GPU PIPELINE

DreamTuber

A checkpointed generation pipeline orchestrating GPU workers to render long-form video from one prompt.

~50% GPU cost cut · Q8_0 GGUF + distilled LTX tier

RunPod serverless GPU · LTX-Video · Chatterbox TTS · GGUF int8 · Cloudflare R2 · Next.js

COMPLETE SYSTEM OVERVIEW
User sends message on Telegram
AGENT · agent.ts
Gather Context
Memory
mem0
+
RAG
vector search
+
History
SQLite
LLM Call
System
+
Memory
+
RAG
+
History
+
Tools
GPT-4o
Tool Routing
Telegram Tools
12
GitHub MCP
26
Notion MCP
21
ReAct Loop · Think → Act → Observe → Repeat
Think
Act
Observe
Repeat
Final Response + Save new memories
Reply sent to user on Telegram
MCP · AGENT RUNTIME

ClawStart

An agent runtime that connects to external MCP servers and gives them durable memory.

47 MCP tools (GitHub 26 + Notion 21) · 12 Telegram tools

MCP · mem0 · ChromaDB · ReAct loop · TypeScript · Docker

LANGGRAPH · STATEGRAPH · 8 NODES
WhatsApp in
text · voice → Whisper v3
Memory extraction
Qdrant · MiniLM-L6
Router
Llama 3.1 8B
Context injection
persona · schedule
Memory injection
Qdrant · top-3
Conversation
Llama 3.3 70B
Image
Pollinations
Audio
ElevenLabs Flash
Summarise
> 20 msgs
WhatsApp out
LANGGRAPH · MULTI-MODAL AGENT

WAMind

A LangGraph state machine running an autonomous multi-modal agent over WhatsApp.

8-node graph · Qdrant long-term + SQLite checkpoints

LangGraph · Qdrant · Groq Llama 3.3 70B · Whisper v3 · ElevenLabs · FastAPI · Docker

LIVE DEMO · 4×
MULTI-LLM ROUTING

ResumeAI Maker

A provider-agnostic routing layer fronting five LLM providers behind one typed interface.

5 interchangeable providers · 52+ ATS templates

Vercel AI SDK v6 · Anthropic/OpenAI/Google/Groq/Ollama · Drizzle/Postgres · better-auth · Puppeteer

04

EXPERIENCE

Professional
experience.

Where I've worked, and what I actually shipped there.

AI Engineering Intern

Logictech Digital

Feb 2026 – Mar 2026

Karachi, Pakistan

  • Built and optimised production RAG pipelines: semantic search, document chunking and embeddings-based retrieval over enterprise knowledge bases.
  • Reduced query latency from ~3s to ~1s through response caching, ANN index tuning and top-k optimisation.
  • Developed LLM orchestration layers and agent-based systems in a cross-functional team and shared codebase, shipping features integrated via REST APIs.
RAGLLM orchestrationSemantic searchREST APIs

EEG Data Analysis & Pattern Recognition Intern

Embedded Neuro Systems Lab · DUET

Jul 2025 – Aug 2025

Karachi, Pakistan

  • Processed and classified EEG signals using ICA, FFT and digital filtering across multi-channel recordings.
  • Engineered preprocessing and feature-extraction pipelines that improved neural-signal classification for the lab's ongoing research.
Signal processingICA / FFTClassificationPython
05

EDUCATION

Where it
started.

BS in Artificial Intelligence

Dawood University of Engineering & Technology (DUET)

2022 – 2026

Karachi, Pakistan

Specialized in applied machine learning and deep learning, with coursework spanning NLP, computer vision and the systems fundamentals behind them.

RELEVANT COURSEWORK

Machine LearningDeep LearningComputer VisionNatural Language ProcessingData Structures & AlgorithmsDatabase SystemsProbability & StatisticsComputer Networks
06

TECH STACK

The stack behind
the systems.

Grouped the way the work actually breaks down, from orchestration through to what keeps it alive in production.

Agents & Orchestration

LangGraphLangChainCrewAIMCPmem0ReActTool callingAgent harness

RAG & Retrieval

QdrantFAISSChromaDBPineconeHybrid search (BM25)RerankingEmbeddings

Eval & Observability

RAGASLangSmithLangfuseLLM-as-judgePrompt versioningTracing

Self-Hosted Inference & Serving

SLMsRunPod GPU workersQuantizationGGUFOllamaModel routingSemantic caching

LLM Foundations

Transformer architectureAttentionSelf-attentionKV cachePagedAttentionGrouped-Query Attention (GQA)

Fine-Tuning & Pre-Training

SLM fine-tuningSLM pre-trainingTokenizer pre-training

Models & Providers

ClaudeGPTGemini 2.5Llama 3.1GroqOllamaHugging Face

Platform & LLMOps

FastAPIDockerKubernetesPrometheusGrafanaJenkinsGitHub ActionsGCP
OpenAI ·Anthropic ·Gemini ·Groq ·SLMs ·Langfuse ·RAGAS ·FAISS ·Qdrant ·ChromaDB ·LiveKit ·Kubernetes ·Prometheus ·Grafana ·OpenAI ·Anthropic ·Gemini ·Groq ·SLMs ·Langfuse ·RAGAS ·FAISS ·Qdrant ·ChromaDB ·LiveKit ·Kubernetes ·Prometheus ·Grafana ·
07

CREDENTIALS

Verified
certifications.

08

LET'S WORK TOGETHER

Have a role or product
in mind? Let's build
something real.

Location

Karachi, Pakistan

Available

For new opportunities