I'm an AI/ML engineer who builds agentic systems, LLM tooling, and retrieval pipelines. I like agents that do real work, retrieval that stays grounded in its sources, and evaluation you can actually trust.
- LegalBench-Agent: an LLM agent built and evaluated on the LegalBench suite for legal reasoning tasks.
- agentic-rag: a retrieval system where an agent plans, retrieves, and reasons over documents instead of doing a single-shot lookup.
- rag-benchmarking: a harness for measuring and comparing RAG pipelines on retrieval quality and answer accuracy.
Python 路 TypeScript 路 React / Next.js 路 Node 路 LangGraph 路 MCP 路 RAG 路 Postgres 路 Redis 路 Supabase



