Shivam Bhardwaj
Fig. 0
AI engineer. Agents, evals, and the interfaces around them.
Fig. —
I build AI systems, and the product around them.
I design and build AI systems end to end: the pipelines, the agents, the evals that keep them honest, and the product people actually use. I started in design, so both halves matter to me.
- 5+
- years shipping production software
- 3
- agent-evaluation tools built in 2026
- 2
- open-source packages, PyPI & npm
Fig. —
An agent is a system, not a prompt.
Ground it in real data.
Retrieval over real data, with sources you can point to.
Make the flow explicit.
LangGraph state machines, not one heroic prompt.
Tools, with a human on the consequential ones.
Agents act; anything irreversible waits for a human.
Evals before demos.
Judges score it, deterministic gates block it, baselines catch regressions.
Stream it, watch it, keep it cheap.
Streamed, observed, and routed to the right model.
A replayed run
› Can we deliver the Northwind spring campaign by Friday?
- parse_requestorchestrateintent: delivery promise · deadline: Fri · scope: campaign
- get_commitmentscontext
- check_capacityact
- check_suppliersact
- read_calendarcontext
- schedule_gateevaluate
- verdictship
- await_approvalact
- monitororchestrate
- supplier_delaycontext
- re_evaluateevaluate
A reconstruction of the Can I Say Yes? demo flow, abridged. Names and timings are illustrative.
Fig. —
Things I built to find out how.
EvalRun
A quality score is not a release decision.
Local-first testing for AI agents: LLM judges, hard policy gates, regression baselines, one CI exit code.
Scree
One field of points, any shape.
A WebGL engine that morphs one field of points between images, words and 3D. This page runs on the same idea.
Prism
Evidence before an agent goes to a customer.
Compares agent releases on a versioned suite and records who approved what.
Can I Say Yes?
Before you promise a customer, let the agent check reality.
An investigator agent answers SAFE, UNSAFE or UNKNOWN with cited evidence, then keeps watching.
DreamFrame
Tell a story, watch it get drawn.
A director agent and five specialists turn a prompt, voice or photo into a streamed graphic novel.
LLM Gateway Lite
Route, execute, evaluate, recover.
Gemini-first routing, a LangGraph pipeline and sampled LLM-judge evals, with an observability UI.
Fig. —
Feb 2025 — PresentRemote
NutriCheck
Senior Full Stack / AI Systems Engineer
- LangGraph pipelines for lab reports, biomarkers and nutrition plans.
- Voice-first health agent with streaming TTS; FastAPI and NestJS services.
Feb 2025 — Jun 2025Remote · Contract
KIWIQ.ai
Founding Frontend Engineer
- Rebuilt an AI content platform and its workflow builder.
- LLM pipelines turning LinkedIn data into content strategy.
Sep 2023 — Dec 2024New Delhi
QPe · Appernity
Full Stack Engineer
- Next.js dashboards for 100K+ monthly visitors.
- Five client products across React, Vue, Ember and Ruby.
Aug 2022 — Aug 2023Noida
DropTheQ
Frontend Developer
- Load time from 3.0s to 2.5s; 4.5/5 user satisfaction.
Dec 2020 — Jul 2022Remote
FlixStock
UX/UI Designer → Frontend Developer
- Three UI overhauls in six months, from Figma to production.
Fig. —
The workbench, roughly by reach.
Applied AI
LangGraph, LangChain, RAG pipelines, Pinecone, FAISS, Weaviate, OpenAI, Gemini, Bedrock, PyTorch, Agent evaluation
Backend
Python, FastAPI, Django, Node.js, NestJS, Go, PostgreSQL, MongoDB, Redis, GraphQL, Kafka
Frontend
TypeScript, React, Next.js, Vue, Nuxt, WebGL, Framer Motion, Zustand, React Native
Infrastructure
Docker, Kubernetes, AWS, GCP, CI/CD, WebSockets, Stripe, FHIR
Fig. —
Building something that has to hold up?
Open to AI engineering roles and projects.