Observability
50Companies categorized as Observability.
Datadog provides a cloud-based monitoring and observability platform that collects metrics, traces, logs, and other data to help developers, IT operations, and security teams monitor applications and infrastructure.
Sentry provides application performance monitoring and error tracking tools that developers and software teams use to capture and debug errors and performance issues in their code.
Elastic provides the Elasticsearch platform, a set of search, observability, security, and AI tools that organizations use to store, retrieve, and analyze their data. It is used by enterprises and developers handling large volumes of information.
New Relic offers an observability platform that collects and correlates telemetry from applications, infrastructure, and AI services, allowing engineers to diagnose issues and reduce mean time to recovery.
Grafana Labs provides an observability platform that aggregates metrics, logs, traces, and other telemetry data into a single cloud service used by developers and site reliability engineers. The platform supports open standards such as OpenTelemetry and integrates with many monitoring tools.
Splunk provides a data platform that lets enterprises collect, analyze, and act on machine-generated data for security, observability, and operational intelligence. It is used by IT, security, and development teams to monitor systems, detect threats, and improve performance.
Honeycomb provides an observability platform that ingests logs, metrics, traces, and other structured telemetry, allowing engineers and AI agents to query and troubleshoot production systems. Engineering teams use it to monitor distributed services and AI workloads.
Braintrust provides an AI observability platform that lets engineering and product teams trace production AI calls, run evaluations, and detect regressions before deployment. It includes real‑time trace inspection, automated pattern discovery, and tools for creating evaluation datasets.
Sumo Logic provides a cloud-based platform for log management, monitoring, and security information and event management (SIEM) that lets developers, IT operations, and security teams collect, analyze, and act on log data from cloud and on-premises systems.
Mezmo provides a telemetry data platform that collects logs, metrics, and traces and uses AI-driven observability to help site reliability engineers and developers detect issues and identify root causes.
Rollbar provides error monitoring and tracking for software teams, offering real-time alerts, session replay, and root cause analysis. Developers use it to detect, investigate, and resolve errors in their applications.
Observe offers an observability platform that stores logs, metrics, traces, and AI application data in an open data lake and uses AI to assist troubleshooting. It is used by developers and site reliability engineers to monitor and investigate system performance.
Langfuse is an open-source AI engineering platform that provides tracing, prompt management, evaluation, and analytics for LLM applications. It enables developers and teams to monitor, debug, and improve their language model applications.
LogRocket provides session replay, analytics, error tracking, and performance monitoring tools for web and mobile applications. It lets developers and product teams view user sessions and diagnose issues affecting their users.
Coralogix offers an AI‑native observability platform that ingests, stores, and queries telemetry data for developers and operations teams. The platform provides a unified query language and integrates with tools such as Slack and PagerDuty.
BugSnag provides error and performance monitoring tools for mobile and web applications. It helps development teams detect, prioritize, and fix bugs and performance issues.
Bindplane offers an OpenTelemetry‑native telemetry pipeline that collects, processes, and routes logs, metrics, and traces for observability and security teams.
Axiom provides an agent-native platform for logs, metrics, traces, and events, allowing engineering teams to collect, store, and analyze event data at scale.
Papertrail is a log management service that lets developers and operations teams collect, search, and monitor logs from their applications in real time. It offers centralized log ingestion, grouping, and alerting features.
Monte Carlo offers a platform for data and AI observability that lets enterprise teams monitor, troubleshoot, and improve production AI systems and agents. It is used by data engineering and analytics teams to gain visibility across data pipelines and AI models.
Loggly provides a cloud-based log management and analysis service that collects, stores, and enables users to search and visualize log data from applications and infrastructure. It is used by developers and operations teams.
Raindrop is a monitoring and observability platform for AI agents that records each run, detects failures, and sends alerts to tools like Slack. It is used by developers and engineering teams building production AI agents.
pganalyze provides a platform for PostgreSQL performance monitoring and tuning that lets database administrators and developers collect query statistics, analyze execution plans, and receive configuration and indexing recommendations.
Airbrake provides error monitoring and performance monitoring for applications, giving developers real-time alerts and insights into errors and performance metrics across multiple programming languages.