Galileo is a comprehensive LLM evaluation platform for data scientists and ML engineers to rapidly test, debug, and monitor prompts, fine-tunes, and RAG systems from development through production.
Galileo is a machine learning operations (MLOps) platform specifically designed for the lifecycle of large language models (LLMs). It provides a suite of tools to help AI teams systematically evaluate, debug, and observe their models to ensure quality and reliability. The platform primarily serves data scientists and machine learning engineers who are building applications on top of LLMs, including retrieval-augmented generation (RAG) systems. Its unique value proposition lies in its focus on the entire development process, offering automated metrics for detecting issues like hallucinations and toxicity, alongside tools for collecting human feedback. By providing real-time production monitoring and 'guardrails' to prevent errors, Galileo aims to accelerate the deployment of trustworthy AI applications.
Data scientists, machine learning engineers, and AI development teams building, deploying, and monitoring applications powered by large language models (LLMs) and retrieval-augmented generation (RAG) systems.
Based on 0 reviews
2021
San Francisco, USA
Community
Free forever for individuals and small teams. Includes all core evaluation workflows for one project and a limited number of evaluations.
Free
Enterprise
For teams building production-grade AI applications. Includes unlimited projects, advanced security features, SSO, role-based access control (RBAC), premium support, and on-premise deployment options.
Contact Sales
Arize AI is a broad ML observability platform, making it a strong choice for teams needing to monitor traditional ML models in addition to their LLM applications within a single tool.
Tightly integrated with the LangChain framework, LangSmith is the ideal choice for developers who are already building extensively within that specific ecosystem.
W&B is a strong alternative for teams who already use their popular platform for experiment tracking and want an integrated solution for LLM evaluation and prompts management.
Fiddler AI offers a comprehensive Model Performance Management platform with deep explainability (XAI) features, appealing to organizations with strict model governance and compliance requirements.
Join thousands of users and see how Galileo can transform your workflow today.
Visit Galileo