What I've built

Deployments, documented

Production systems delivered in regulated environments — pharma, audit, and enterprise. Each with the metric that mattered.

PHARMA · MEDICAL WRITINGIn Production

Agentic ICF Generator

Transforms complex Clinical Study Protocol tables spanning multiple pages into concise ICF Summary Tables using Skills Cards architecture with scoped MCP tool restrictions.

DSPyMCPSkills CardsVertex AI
<5 min
from 1–2 days manual
PHARMA · DOCUMENT AIIn Production

Vision Agentic RAG

Multimodal retrieval system with custom Docling parser extracting tables, figures, and images from PDFs, RTF, DOCX into structured Markdown. Indexed into Weaviate and GCS.

DSPyWeaviateDoclingGCSColPali
60→85%
retrieval hit-rate@6
PHARMA · LLM EVALUATIONIn Production

LLM-Judge Evaluation Stack

Claim-level evaluation for AI-generated clinical study report sections: 4 judges (faithfulness, completeness, patient-safety, alignment) issuing boolean verdicts with source-snippet citations, abstention-aware metrics, MLflow-tracked runs and DVC-versioned golden datasets on Vertex AI. Judges meta-evaluated via seeded-error datasets, calibration against medical-writer annotations (κ + AC1), and cross-model verdict verification (Gemini vs GPT).

DSPyLLM-as-Judgeκ/AC1MLFlowDVCVertex AI
1.00
grounding · completeness +7 pts
AUDIT · ENTERPRISE RAGDelivered

Compound AI Audit Chatbot

RAG chatbot for KPMG's audit department parsing thousands of documents (PDF, PPTX, images) with query decomposition, chain-of-thought reasoning, and multi-hop traversal.

DSPy OptimizersAzure SearchLangFuseStreaming
70→94%
answer accuracy