Future AGI Cookbooks: Guides and Tutorials
Runnable recipes for evaluation, tracing, simulation, and optimization, organized by what you are building and by platform feature.
Start with a quickstart, or jump straight to the recipes for what you’re building or the platform feature you need.
Start Here
Your First Evaluation
Score LLM outputs for hallucination, toxicity, and custom criteria
Dataset Management
Create, edit, and evaluate a dataset from the dashboard
Dataset SDK Batch Eval
Upload a CSV, run batch evals, and download scored results
Evaluator SDK Basics
Run built-in eval templates with the ai-evaluation package
Self-Hosted Docker Compose
Deploy the full open-source stack locally in five minutes
Monitoring and Alerts
Track latency and cost trends, then set threshold alerts
By Use Case
Chat & Support Agents
Evaluate and simulate conversational agents, 7 recipes
RAG & Document Q&A
Score retrieval, generation, and grounding, 9 recipes
Voice Agents
Test voice agents with scripted call simulations, 2 recipes
Multi-Agent & Tool Use
Instrument and evaluate multi-agent and tool-calling systems, 5 recipes
Content & Multimodal
Score text, image, audio, and PDF outputs, 5 recipes
Text-to-SQL
Build and evaluate text-to-SQL agents, 2 recipes
By Platform Feature
Questions & Discussion