Galileo
AI observability and evaluation platform that turns offline evals into production guardrails for LLM and agent apps.
What it does
Galileo is an AI observability and evaluation platform for teams building LLM and agent applications. It helps build datasets and custom evaluators, auto-tunes metrics from live feedback, and distills expensive LLM-as-judge evals into compact 'Luna' models that monitor production traffic at lower cost. It also surfaces failure modes and prescriptive fixes to speed debugging.
Core features
20+ out-of-the-box evals for RAG, agents, safety and security
Custom evaluator building
Auto-tuned metrics from live feedback
Luna models for low-cost production guardrails
Failure-mode detection with fix suggestions
Dataset and groundtruth management
Best for
→Evaluate LLM and agent app quality
→Monitor AI systems in production
→Catch hallucinations and unsafe outputs
→Debug and improve agent failures
Toolspool rankingby global site rank
Reviews
Big-picture takes: what it's for and whether it delivers. High-engagement YouTube videos — not sponsored.
Tutorials
Step-by-step: exactly how to get things done with it.