I build practical generative AI solutions and equally focus on evaluating whether those systems are reliable, safe, and useful in the real world.
Current Focus:
- ๐งช LLM Evaluation & Responsible AI - Testing safety, hallucination, bias, guardrails
- ๐๏ธ Voice & Agentic AI - Building and evaluating voice agents
- ๐ Agentic Systems - Long-running, autonomous AI agents
- ๐ Multilingual/Indic AI - Voice, evaluation, synthetic data for Indian languages
AI Evaluation & Testing Platform for Voice & Agentic AI
An automated platform that:
- ๐ญ Generates realistic personas and multi-turn scenarios
- ๐ค Interacts with AI agents to test them rigorously
- ๐ก๏ธ Evaluates: safety, hallucination, instruction-following, bias, guardrails
- ๐ Produces evidence-based evaluation reports
- ๐ Uses LLM-as-a-Judge for intelligent evaluation
- GenAI/LLM architecture and optimization
- Evaluation frameworks & metrics
- Responsible AI & safety testing
- Python & experimentation
- Building realistic test scenarios
AI/ML: Claude, GPT, LangChain, AWS Bedrock, RAG, Vector DBs
Voice: Speech Recognition, TTS, Voice Cloning
Data: Synthetic Data Generation, Multilingual NLP
Backend: Python, FastAPI, Jupyter
I'm seeking builders and experimenters with complementary strengths:
- ๐๏ธ Full-Stack/Product Development - Build UI, infrastructure, deployment
- ๐ค AI/ML & Agents - Implement evaluation logic, agent orchestration
- ๐๏ธ Speech/Voice AI - Voice evaluation, TTS, speech processing
- ๐จ UI/UX Design - Make evaluation reports beautiful and actionable
What I value:
- โ Shipping code over endless discussions
- โ Iterating quickly and learning fast
- โ Challenging assumptions
- โ Curiosity and experimentation
Real-time voice AI agent evaluation framework. Tests guardrails, compliance, safety, and edge cases.
Advanced RAG system with multi-turn reasoning and evaluation capabilities.
Integration utilities for AWS Bedrock LLMs.
ACTIVELY SEEKING TEAMMATES FOR AI EVALUATION PLATFORM
If you're interested in building an AI evaluation platform and have complementary skills, let's connect!
- Open to collaboration on AI evaluation, voice agents, Indic AI
- Excited about quick iterations and shipping working prototypes
- Looking for curious builders, not idea discussers
"The best way to predict the future is to build it." โ Let's build something great together! ๐

