LLMOps - Top Apps & Software
LLMOps platforms manage prompts, evaluation, fine-tuning, monitoring, and cost for large language model applications in production.

Dify
Dify is an open-source platform to build, run, and monitor LLM-powered apps using visual workflows, RAG pipelines, agents, model management, APIs, and self-hosting options.

LangSmith
LangSmith is a platform for developing and optimizing LLM applications, offering tools for tracing, monitoring, evaluation, and debugging throughout the application lifecycle.

Wordware
Wordware is an IDE for collaboratively developing and deploying AI agents and applications, featuring tools for prompt management, API integration, and workflow optimization.

Langfuse
Langfuse is an open-source platform for managing and evaluating AI applications, enabling teams to monitor, debug, and optimize large language models effectively.

Klu.ai
Klu.ai is a platform for designing, deploying, and optimizing AI applications, focusing on collaborative prompt engineering and multi-LLM integration.

Humanloop
Humanloop is a platform for developing and optimizing LLM applications with human feedback and real-time monitoring, enhancing model accuracy and adaptability.

Agenta
Open-source platform to build, test, monitor, and evaluate LLM applications and agents, with prompt versioning, A/B testing, cost/latency observability, tracing, and feedback integration.

LLM Gateway
LLM Gateway provides a unified API for managing and analyzing LLM requests across multiple providers, with features for usage tracking and access controls.

Laminar AI
Laminar AI is an open-source platform for monitoring and analyzing Large Language Model applications, offering tools for observability, analytics, and prompt management.

Promptly
Promptly is a low-code platform for enterprises that simplifies creating, testing, and managing prompts for large language models.

Arkor
Arkor helps developers prepare datasets, train open-weight models on managed GPUs, evaluate them, and deploy them through OpenAI-compatible APIs.

Promptmonitor
Promptmonitor allows marketers to track and optimize their brand's visibility across various AI platforms to enhance traffic, lead generation, and sales.

Inworld
Inworld provides an AI runtime for building and scaling consumer apps: automated MLOps, live A/B experiments, real-time voice (TTS) support, and flexible deployment for production use.

Traceloop
Traceloop monitors and tests Large Language Model applications, providing insights, alerts, and tools for optimizing performance and debugging workflows.

Vellum
Vellum is a development platform for creating, deploying, and monitoring AI applications, supporting multiple LLM providers and promoting collaboration among teams.

Helicone
Helicone is an open-source framework for monitoring and optimizing large language models, supporting workflow tracking, prompt experimentation, and user segmentation.

Galileo
Platform for evaluating, monitoring, and securing generative AI apps, with tools for testing prompts, tracing outputs, and adding production guardrails.

Maxim AI
Maxim AI is an evaluation and observability platform for AI teams to develop and deploy AI applications with tools for testing, monitoring, and data management.

Confident AI
Confident AI is a platform for evaluating and monitoring LLM applications, offering tools for testing, benchmarking, and improving model performance with structured workflows.

Mastra
Mastra is an open-source TypeScript framework to build and deploy AI applications—agents, durable workflows, RAG memory, tool integration, evals, and cloud deployment with observability.

Lamini
Lamini is an enterprise platform that enables teams to develop and manage customized large language models using proprietary data in secure environments.

GradientJ
GradientJ helps teams deploy and manage large language models, offering tools for performance tracking, prompt building, and continuous model improvement.

AI Monitor
Monitors how brands' content appears and performs on AI-generated platforms (e.g., ChatGPT, Google AI Overview) and reports data to track and improve visibility.

TrueFoundry
TrueFoundry is a PaaS for machine learning teams to build, deploy, and manage AI applications efficiently on their own cloud or on-premise infrastructure.

Test AI Models
Run your prompts across multiple AI models at once and view real-time comparisons of output quality, speed, and cost without setup or API keys.

Deepchecks
Deepchecks simplifies the evaluation of large language models and machine learning systems, providing automated checks, bias detection, and performance monitoring.

Orq.ai
Orq.ai is a platform for AI teams to develop, deploy, and optimize applications using large language models, with tools for integration, testing, monitoring, and compliance.

FinetuneDB
FinetuneDB is a platform that helps users create, manage datasets, and fine-tune large language models for improved performance and customization.

Weavel
Weavel is a platform for performance analytics and testing of LLM applications, allowing users to create datasets, conduct tests, and monitor usage.

Patronus AI
Patronus AI is an automated platform that tests and monitors Large Language Models, helping enterprises ensure accuracy and reliability in generative AI applications.

SiliconFlow
Cloud API for developers to run, fine‑tune, and deploy language and multimodal models with low‑latency inference, scalable deployment options, SDKs, observability, and no user data retention.

Braintrust
Braintrust is a platform for developing and evaluating AI applications, offering tools for model optimization, collaboration, and dataset management.

CoreCtic AI
AI proxy and analytics platform that connects multiple language model providers, routes requests, reduces costs with semantic caching, and monitors usage, performance, and client billing.

Not Diamond
Not Diamond routes prompts across multiple LLMs, adapts prompts for different models, and predicts the best model per input to improve accuracy, cost, and reliability.

LLMTest
Proxies OpenAI and Anthropic calls, tracks costs, benchmarks 340+ models, and optimizes prompts and routing using live traffic.

Entry Point AI
Entry Point AI is a platform for managing, training, and evaluating large language models, allowing users to fine-tune model performance and collaborate on training tasks.

Augento
Augento is a fine-tuning service for AI models, allowing users to improve agent performance by providing feedback for better accuracy and consistency.

Ultra AI
Ultra AI is a centralized platform for developers to access multiple AI providers, manage prompts, improve performance through caching, and analyze API usage.

AgentOps
AgentOps is a developer platform for testing and debugging AI agents, offering tools to record prompts, completions, and monitor agent performance.

Fiddler
AI observability and security platform for monitoring, explaining, and governing ML models, LLM apps, and agentic AI systems across development and production.

MosaicML
MosaicML is a platform for building, training, and deploying domain-specific AI models, focusing on data security, model governance, and supporting complex AI workflows.

Gentrace
Gentrace is a developer platform for testing and monitoring AI applications, enhancing efficiency and reliability in development and deployment processes.

AnyApiAI
A unified API platform for accessing 400+ AI models, with routing, retries, analytics, and tools to test and manage AI app integrations.

Bellaire AI
Bellaire AI is a platform for building and deploying AI agents and chatbots using large language models and automation tools for businesses.

Parea
Parea AI assists developers in rapidly creating and managing production-ready LLM products, allowing for prompt experimentation, versioning, and performance evaluation.

Relyable
Relyable runs simulated conversations between AI voice agents, evaluates live calls, and integrates with Vapi and Retell AI.

Toolhouse
Toolhouse is a Backend-as-a-Service platform that helps developers create, deploy, and manage AI agents as APIs with minimal coding.

LaminarFlow
LaminarFlow automates content creation and publication, streamlining workflows by managing tasks, integrating data sources, and ensuring accuracy.

Pioneer
Platform for fine-tuning open-source language models with automated data generation, training, evaluation, deployment, and ongoing improvement from live inference data.

Lang.ai EU
Lang.ai is a no-code platform that allows customer support teams to automate tasks and manage AI models, integrating with Zendesk and Salesforce.

Athina AI
Athina enables developers to monitor and assess their LLM applications with visibility into RAG pipelines and over 40 preset evaluation metrics.

Mosaic
Mosaic is an AI video editing app that allows users to describe edits in natural language, making the editing process much faster and more efficient.

PromptPilot
A platform for building, testing, and improving prompts for large language model apps, with API-based feedback and iterative optimization.

Parlance
Parlance uses conversational AI to enhance phone communication for businesses, allowing users to connect easily and naturally while providing managed services for success.

OrcaRouter
OpenAI-compatible API gateway that routes prompts across 200+ models from multiple providers, with fallback, cost tracking, and one unified endpoint.

Tracer
Tracer routes requests among AI models, combining their outputs and analyzing performance to deliver efficient, task-specific results through an OpenAI-compatible API.

Nirixa AI
An AI observability platform that tracks LLM calls, token costs, prompt drift, hallucination risk, and latency across providers in real time.

ModelRiver
A unified API gateway for multiple LLM providers offering streaming responses, configurable settings, routing, rate limits, failover, built-in packages, and usage analytics.

Draft'n run
Open-source platform to design, deploy, and monitor AI agents and automated workflows with built-in DevOps, observability, and production integrations.

Flutch
Flutch deploys and tests AI agents with observability, tracing, quality checks, cost controls, versioning and multi-channel deployment via CLI; supports no-code setup and framework-agnostic integrations.