# Changelog

New features, improvements, and fixes across the Openlayer platform, release by release.

- [AI summaries, semantic search filters, and the remote MCP connector](/changelog/summer-2026.md): Sessions, traces, and test results now open with an AI summary — how the session went, what a trace did, and the dominant failure modes behind a test result, with drill-down from each failure mode to the exact rows behin...

- [Trace timeline, project lifecycles, and Claude Agent SDK tracing](/changelog/june-2026.md): A new Gantt-style view of any trace. Every span lines up by start time, duration, and nesting, so the slow step or the unexpected branch is obvious at a glance — no more reading a flat list to reconstruct what ran when.

- [Openlayer Gateway, customizable dashboards, and Salesforce Agentforce GA](/changelog/spring-2026.md): Openlayer Gateway is a new managed control layer for every LLM call your stack makes. The admin portal handles API-key management, per-key rate limits and freeze, and team-based organization — and every request that flow...

- [Stronger security, multimodal tracing, and expanded governance controls](/changelog/stronger-security-multimodal-tracing-and-expanded-governance-controls.md): This release strengthens enterprise-grade AI governance with platform-wide multi-factor authentication (MFA), deployment-wide directory sync, and more granular project-level access controls. We’ve expanded observability...

- [Batch test re-runs and expanded integrations](/changelog/batch-test-re-runs-and-expanded-integrations.md): We now support re-running tests in batch on Openlayer, so that you can recalculate results when new data is added retroactively. In addition, we’ve added a ton of new integrations – Ruby and Strands Agents SDKs, Gemini A...

- [Pausing tests, checks for duplicate keys, and new integrations](/changelog/pausing-tests-checks-for-duplicate-keys-and-new-integrations.md): This month, we shipped a range of improvements across Openlayer, including new integrations, tests, and developer features. We’re also now available on the AWS, Azure, and Google Cloud marketplaces, and we’ve added suppo...

- [Introducing Openlayer Governance](/changelog/introducing-openlayer-governance.md): Introducing Openlayer Governance: the fastest way to track and enforce rules for your AI systems. With custom frameworks, you can define rules like:

- [ Support for tracking users and sessions, test tags, and more](/changelog/support-for-tracking-users-and-sessions-test-tags-and-more.md): In addition to traces, you can now track users and sessions in the Openlayer app. See the full session and record history for each user to better understand your system behavior. Head to the Data section in Openlayer tod...

- [Test bundles, new tests, support for new Python runtimes](/changelog/test-bundles-new-tests-support-for-new-python-runtimes.md): We’re very excited to introduce test bundles to the Openlayer platform! Easily create a set of tests related to use cases or policies of interest, such as the EU AI Act, OWASP, agentic workflows, data quality, and more....

- [Complete design system overhaul, Snowflake Integration](/changelog/complete-design-system-overhaul-snowflake-integration.md): This month, we’re excited to unveil our brand new UI! We’ve defined an improved design system, including updated and thoughtfully-crafted styles and components to give the product a fresh, engaging look and feel. The new...

- [The Openlayer MCP server, Automatic thresholds, BigQuery Integration and Anomaly Detection, Project-level access groups](/changelog/the-openlayer-mcp-server-automatic-thresholds-bigquery-integration-and-anomaly-detection-project.md): We’re introducing an exciting new feature to our observability platform: automatic thresholds for tests and anomaly detection.

- [Project-level secrets, tracing LLM requests with OpenTelemetry](/changelog/project-level-secrets-tracing-llm-requests-with-opentelemetry.md): We’ve shipped new ways to manage secrets and API keys across your Openlayer projects, making it easier to scale and stay secure.

- [SAML Directory Sync, new LLM-as-a-judge models, and website refresh](/changelog/saml-directory-sync-new-llm-as-a-judge-models-and-website-refresh.md): We’ve added lots of features and enhancements across our platform, focused on improving performance, expanding functionality, and streamlining workflows. To highlight a few:

- [Improved test diagnosis page, SAML SSO, design refreshes, + more](/changelog/improved-test-diagnosis-page-saml-sso-design-refreshes-more.md): 🔎🩹 Quickly identify issues with the improved test diagnosis page Diagnosing issues is a core part of the eval process, and that’s why we want to make sure our test diagnosis page is as helpful as possible. To make it e...

- [Custom metrics, rotating API keys, and new models for direct-to-API calls](/changelog/custom-metrics-rotating-api-keys-and-new-models-for-direct-to-api-calls.md): We understand that you may have metrics that are highly specific to your use case, and you want to use these alongside standard metrics to eval your AI systems. That’s why we built custom metrics. You can now upload any...

- [Improved quality control over your LLM’s responses with annotations and human feedback](/changelog/improved-quality-control-over-your-llm-s-responses-with-annotations-and-human-feedback.md): Setting up alerts is an essential first step to monitoring your LLMs, but in order to understand why issues arise in production, it’s helpful to have human eyes to review requests.

- [ Simple, dev-focused workflow for AI evals](/changelog/simple-dev-focused-workflow-for-ai-evals.md): Most of us get how crucial AI evals are now. The thing is, almost all the eval platforms we’ve seen are clunky – there’s too much manual setup and adaptation needed, which breaks developers’ workflows.

- [Trace every step of your requests](/changelog/trace-every-step-of-your-requests.md): We’re thrilled to share with you the latest update to Openlayer: comprehensive tracing capabilities and enhanced request streaming with function calling support.

- [More tests around latency metrics](/changelog/more-tests-around-latency-metrics.md): We’ve added more ways to test latency. Beyond just mean, max, and total, you can now make test latency with minimum, median, 90th percentile, and 99th percentile metrics. Just head over to the Performance page and the ne...

- [Go deep on test result history and add multiple criteria to GPT evaluation tests](/changelog/go-deep-on-test-result-history-and-add-multiple-criteria-to-gpt-evaluation-tests.md): You can now click on any test to dive deep into the test result history. Select specific date ranges to see the requests from that time period, scrub through the graph to spot patterns over time, and get a full picture o...

- [Cost-per-request, new tests, subpopulation support for data tests, and more precise row filtering](/changelog/cost-per-request-new-tests-subpopulation-support-for-data-tests-and-more-precise-row-filtering.md): We’re excited to introduce the newest set of tests to hit Openlayer! Make sure column averages fall within a certain range with the Column average test. Ensure that your outputs contain specific keywords per request with...

- [Log multi-turn interactions, sort and filter production requests, and token usage and latency graphs](/changelog/log-multi-turn-interactions-sort-and-filter-production-requests-and-token-usage-and-latency.md): Introducing support for multi-turn interactions. You can now log and refer back to the full chat history of each of your production requests in Openlayer. Sort by timestamp, token usage, or latency to dig deeper into you...

- [GPT evaluation, Great Expectations, real-time streaming, TypeScript support, and new docs](/changelog/gpt-evaluation-great-expectations-real-time-streaming-typescript-support-and-new-docs.md): Openlayer now offers built-in GPT evaluation for your model outputs. You can write descriptive evaluations like “Make sure the outputs do not contain profanity,” and we will use an LLM to grade your agent or model given...

- [Enhanced onboarding, redesigned navigation, and new goals](/changelog/enhanced-onboarding-redesigned-navigation-and-new-goals.md): We’re thrilled to announce a new and improved onboarding flow, designed to make your start with us even smoother. We’ve also completely redesigned the app navigation, making it more intuitive than ever.

- [ Evals for LLMs, real-time monitoring, Slack notifications and so much more!](/changelog/evals-for-llms-real-time-monitoring-slack-notifications-and-so-much-more.md): It’s been a couple of months since we posted our last update, but not without good reason! Our team has been cranking away at our two most requested features: support for LLMs and real-time monitoring / observability. We...

- [Regression projects, toasts, and artifact retrieval](/changelog/regression-projects-toasts-and-artifact-retrieval.md): This week we shipped a huge set of features and improvements, including our solution for regression projects!

- [Sign in with Google, sample projects, mentions and more!](/changelog/sign-in-with-google-sample-projects-mentions-and-more.md): We are thrilled to release the first edition of our company’s changelog, marking an exciting new chapter in our journey. We strive for transparency and constant improvement, and this changelog will serve as a comprehensi...
