Changelog
Trace timeline, project lifecycles, and Claude Agent SDK tracing

A new Gantt-style view of any trace. Every span lines up by start time, duration, and nesting, so the slow step or the unexpected branch is obvious at a glance — no more reading a flat list to reconstruct what ran when.
Learn MoreOpenlayer Gateway, customizable dashboards, and Salesforce Agentforce GA

Openlayer Gateway is a new managed control layer for every LLM call your stack makes. The admin portal handles API-key management, per-key rate limits and freeze, and team-based organization — and every request that flow...
Learn MoreStronger security, multimodal tracing, and expanded governance controls

This release strengthens enterprise-grade AI governance with platform-wide multi-factor authentication (MFA), deployment-wide directory sync, and more granular project-level access controls. We’ve expanded observability...
Learn MoreBatch test re-runs and expanded integrations

We now support re-running tests in batch on Openlayer, so that you can recalculate results when new data is added retroactively. In addition, we’ve added a ton of new integrations – Ruby and Strands Agents SDKs, Gemini A...
Learn MorePausing tests, checks for duplicate keys, and new integrations

This month, we shipped a range of improvements across Openlayer, including new integrations, tests, and developer features. We’re also now available on the AWS, Azure, and Google Cloud marketplaces, and we’ve added suppo...
Learn MoreIntroducing Openlayer Governance

Introducing Openlayer Governance: the fastest way to track and enforce rules for your AI systems. With custom frameworks, you can define rules like:
Learn MoreSupport for tracking users and sessions, test tags, and more

In addition to traces, you can now track users and sessions in the Openlayer app. See the full session and record history for each user to better understand your system behavior. Head to the Data section in Openlayer tod...
Learn MoreTest bundles, new tests, support for new Python runtimes

We’re very excited to introduce test bundles to the Openlayer platform! Easily create a set of tests related to use cases or policies of interest, such as the EU AI Act, OWASP, agentic workflows, data quality, and more....
Learn MoreComplete design system overhaul, Snowflake Integration

This month, we’re excited to unveil our brand new UI! We’ve defined an improved design system, including updated and thoughtfully-crafted styles and components to give the product a fresh, engaging look and feel. The new...
Learn MoreThe Openlayer MCP server, Automatic thresholds, BigQuery Integration and Anomaly Detection, Project-level access groups

We’re introducing an exciting new feature to our observability platform: automatic thresholds for tests and anomaly detection.
Learn MoreProject-level secrets, tracing LLM requests with OpenTelemetry

We’ve shipped new ways to manage secrets and API keys across your Openlayer projects, making it easier to scale and stay secure.
Learn MoreSAML Directory Sync, new LLM-as-a-judge models, and website refresh

We’ve added lots of features and enhancements across our platform, focused on improving performance, expanding functionality, and streamlining workflows. To highlight a few:
Learn MoreImproved test diagnosis page, SAML SSO, design refreshes, + more

🔎🩹 Quickly identify issues with the improved test diagnosis page Diagnosing issues is a core part of the eval process, and that’s why we want to make sure our test diagnosis page is as helpful as possible. To make it e...
Learn MoreCustom metrics, rotating API keys, and new models for direct-to-API calls

We understand that you may have metrics that are highly specific to your use case, and you want to use these alongside standard metrics to eval your AI systems. That’s why we built custom metrics. You can now upload any...
Learn MoreImproved quality control over your LLM’s responses with annotations and human feedback

Setting up alerts is an essential first step to monitoring your LLMs, but in order to understand why issues arise in production, it’s helpful to have human eyes to review requests.
Learn MoreSimple, dev-focused workflow for AI evals

Most of us get how crucial AI evals are now. The thing is, almost all the eval platforms we’ve seen are clunky – there’s too much manual setup and adaptation needed, which breaks developers’ workflows.
Learn MoreTrace every step of your requests

We’re thrilled to share with you the latest update to Openlayer: comprehensive tracing capabilities and enhanced request streaming with function calling support.
Learn MoreMore tests around latency metrics

We’ve added more ways to test latency. Beyond just mean, max, and total, you can now make test latency with minimum, median, 90th percentile, and 99th percentile metrics. Just head over to the Performance page and the ne...
Learn MoreGo deep on test result history and add multiple criteria to GPT evaluation tests

You can now click on any test to dive deep into the test result history. Select specific date ranges to see the requests from that time period, scrub through the graph to spot patterns over time, and get a full picture o...
Learn MoreCost-per-request, new tests, subpopulation support for data tests, and more precise row filtering

We’re excited to introduce the newest set of tests to hit Openlayer! Make sure column averages fall within a certain range with the Column average test. Ensure that your outputs contain specific keywords per request with...
Learn MoreLog multi-turn interactions, sort and filter production requests, and token usage and latency graphs

Introducing support for multi-turn interactions. You can now log and refer back to the full chat history of each of your production requests in Openlayer. Sort by timestamp, token usage, or latency to dig deeper into you...
Learn MoreGPT evaluation, Great Expectations, real-time streaming, TypeScript support, and new docs

Openlayer now offers built-in GPT evaluation for your model outputs. You can write descriptive evaluations like “Make sure the outputs do not contain profanity,” and we will use an LLM to grade your agent or model given...
Learn MoreEnhanced onboarding, redesigned navigation, and new goals

We’re thrilled to announce a new and improved onboarding flow, designed to make your start with us even smoother. We’ve also completely redesigned the app navigation, making it more intuitive than ever.
Learn MoreEvals for LLMs, real-time monitoring, Slack notifications and so much more!

It’s been a couple of months since we posted our last update, but not without good reason! Our team has been cranking away at our two most requested features: support for LLMs and real-time monitoring / observability. We...
Learn MoreRegression projects, toasts, and artifact retrieval

This week we shipped a huge set of features and improvements, including our solution for regression projects!
Learn MoreSign in with Google, sample projects, mentions and more!

We are thrilled to release the first edition of our company’s changelog, marking an exciting new chapter in our journey. We strive for transparency and constant improvement, and this changelog will serve as a comprehensi...
Learn More
