Skip to content

Release v0.42.0 - #6207

Merged
samuv merged 1 commit into
mainfrom
release/v0.42.0
Aug 5, 2026
Merged

Release v0.42.0#6207
samuv merged 1 commit into
mainfrom
release/v0.42.0

Conversation

@toolhive-release-app

Copy link
Copy Markdown
Contributor

Release v0.42.0

Version Bump

minor release

Files Updated

  • VERSION
  • deploy/charts/operator-crds/Chart.yaml (path: version)
  • deploy/charts/operator-crds/Chart.yaml (path: appVersion)
  • deploy/charts/operator/Chart.yaml (path: version)
  • deploy/charts/operator/Chart.yaml (path: appVersion)
  • deploy/charts/operator/values.yaml (path: operator.image)
  • deploy/charts/operator/values.yaml (path: operator.toolhiveRunnerImage)
  • deploy/charts/operator/values.yaml (path: operator.vmcpImage)
  • Helm chart docs (via helm-docs)

Next Steps

  1. Review this PR
  2. Merge to main
  3. Release automation will handle the rest

Checklist

  • Version bump is correct
  • All CI checks pass

Release-Triggered-By: samuv
@github-actions github-actions Bot added the size/XS Extra small PR: < 100 lines changed label Aug 5, 2026
@codecov

codecov Bot commented Aug 5, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 72.51%. Comparing base (227ee7b) to head (6103e3b).

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #6207      +/-   ##
==========================================
+ Coverage   72.46%   72.51%   +0.05%     
==========================================
  Files         739      739              
  Lines       76719    76719              
==========================================
+ Hits        55594    55633      +39     
+ Misses      17160    17104      -56     
- Partials     3965     3982      +17     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@samuv
samuv merged commit 2c623d5 into main Aug 5, 2026
45 checks passed
@samuv
samuv deleted the release/v0.42.0 branch August 5, 2026 13:25
@github-actions

github-actions Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

📝 Generated release notes for v0.42.0

Auto-generated by the release-notes skill. Review and, if good, apply with:

gh release edit v0.42.0 --notes-file <paste-below>.md
Click to expand release notes

🚀 Toolhive v0.42.0 is live!

AI-tool plugin management goes end to end — thv ai-plugin gains a full CLI, REST API, and registry catalog — and the skills supply chain gets Sigstore signature verification at install, sync, and upgrade time. Alongside that, a large batch of MCP dual-era correctness fixes lands: multiple clients can finally share a stdio server, and vMCP stops flapping between the Modern and Legacy revisions.

⚠️ Breaking Changes

  • Config CRD status fields removedstatus.referencingWorkloads and status.referenceCount (and the References printer column) are gone from all six config CRDs; replace any automation reading them with a workload field query (migration guide)
  • Cedar policy is now evaluated against the mutated MCP request — if you run a mutating webhook together with authorization, policy decisions and audit records can change on upgrade; re-audit your policies against the post-mutation shape first (migration guide)
  • Recovered HTTP panics no longer produce a log line — an unintended regression from the recovery-middleware migration; without Sentry configured a recovered panic is now silent apart from the 500 (migration guide)
  • Go API removals for out-of-tree importerspkg/telemetry/providers was deleted and two long-published optimizerdec constants were removed (migration guide)
Migration guide: Config CRD status fields removed

Who is affected: anyone reading status.referencingWorkloads or status.referenceCount from MCPOIDCConfig, MCPAuthzConfig, MCPExternalAuthConfig, MCPToolConfig, MCPWebhookConfig, or MCPTelemetryConfigkubectl users relying on the REFERENCES column, scripts and GitOps assertions using jsonpath/jq on those paths, Chainsaw/kuttl tests, kube-state-metrics custom-resource-state configs and the dashboards built on them, and Go code reading .Status.ReferencingWorkloads / .Status.ReferenceCount.

MCPWebhookConfig and MCPTelemetryConfig only ever had referencingWorkloads. MCPTelemetryConfig never had a References printer column, so its kubectl get output is unchanged.

Upgrade safety: these were derived values computed from workload specs — the source of truth (spec.*ConfigRef on workloads) is untouched, so nothing unrecoverable is lost. Applying the new schema does not rewrite or reject existing stored objects; residual values stay inert in etcd until each object's status is next written. No storage-version bump, no CRD delete/recreate, no migration job. Deletion protection is unchanged — every config controller still recomputes referrers live at deletion time and sets DeletionBlocked=True with reason ReferencedByWorkloads.

Before

$ kubectl -n toolhive-system get mcpoidcconfig
NAME       SOURCE   VALID   REFERENCES   AGE
my-oidc    inline   True    3            5d

After

$ kubectl -n toolhive-system get mcpoidcconfig
NAME       SOURCE   VALID   AGE
my-oidc    inline   True    5d

To list referrers, query the workloads by their config-ref:

kubectl -n toolhive-system get mcpservers,mcpremoteproxies,virtualmcpservers -o json \
  | jq -r --arg n my-oidc '.items[]
      | select((.spec.oidcConfigRef.name // .spec.incomingAuth.oidcConfigRef.name) == $n)
      | "\(.kind)/\(.metadata.name)"'

The reference paths per config kind, exactly as the operator's own indexers define them:

Config kind Workload kinds tracked Spec paths
MCPOIDCConfig MCPServer, MCPRemoteProxy, VirtualMCPServer spec.oidcConfigRef.name; spec.incomingAuth.oidcConfigRef.name (vMCP)
MCPAuthzConfig MCPServer, MCPRemoteProxy, VirtualMCPServer spec.authzConfigRef.name; spec.incomingAuth.authzConfigRef.name (vMCP)
MCPTelemetryConfig MCPServer, MCPRemoteProxy, VirtualMCPServer spec.telemetryConfigRef.name
MCPExternalAuthConfig MCPServer, MCPRemoteProxy spec.externalAuthConfigRef.name, or spec.authServerRef.name when spec.authServerRef.kind == "MCPExternalAuthConfig"
MCPToolConfig MCPServer spec.toolConfigRef.name
MCPWebhookConfig MCPServer spec.webhookConfigRef.name

Note kubectl --field-selector will not work for these paths — the operator's indexes are controller-runtime cache indexes, not API-server field selectors. Use -o json | jq or -o custom-columns.

Migration steps

  1. While still on v0.41.x, snapshot anything you may need: kubectl get mcpoidcconfigs,mcpauthzconfigs,mcpexternalauthconfigs,mcptoolconfigs,mcpwebhookconfigs,mcptelemetryconfigs -A -o json > /tmp/thv-config-refs-pre-0.42.json
  2. Grep your automation for referenceCount, referencingWorkloads, and the References/REFERENCES column — shell scripts, kubectl wait --for=jsonpath=, Chainsaw/kuttl assertions, Argo CD/Flux health checks, kube-state-metrics configs, Grafana panels, Kyverno/Gatekeeper rules.
  3. Rewrite each hit with the query for that config kind from the table above. For "is this config still in use?" checks, prefer the condition: kubectl -n NS get mcpoidcconfig my-oidc -o jsonpath='{.status.conditions[?(@.type=="DeletionBlocked")].message}'
  4. helm upgrade the operator-crds chart, then the operator chart. No pre/post hooks needed.
  5. Verify: kubectl -n toolhive-system get mcpoidcconfig shows NAME SOURCE VALID AGE, and deletion of a referenced config still leaves it with DeletionBlocked=True.
  6. Go consumers: drop .Status.ReferencingWorkloads / .Status.ReferenceCount reads. The WorkloadReference type (Kind, Name) is still exported if you want to keep your own list shape.

PR: #5631 — completes the cleanup tracked in #5607

Migration guide: Cedar policy now sees the post-mutation request

Who is affected: only workloads configured with at least one mutating webhook and either Cedar authorization or any consumer of audit / telemetry / usage metrics. Both are shipped, supported, non-mutually-exclusive configurations — thv run --webhook-config <file with a mutating: entry> --authz-config <file>, or MCPWebhookConfig.spec.mutating in the operator. Workloads with no mutating webhook see zero change; the republish is gated on the body actually having changed.

What was wrong: ParsingMiddleware parses the request body once and refuses to parse again. The mutating webhook replaced r.Body but passed the request through unchanged, so Cedar evaluated policy against the tool name and arguments that arrived while the backend executed the ones that ran. The audit half was reachable in the default configuration: the event type and target.name resolve through the parsed-request holder regardless of includeRequestData (which defaults to false), so the audit trail named a request that never executed. Telemetry and usage metrics drifted the same way.

Security framing, stated precisely: before v0.42.0, a client could reach a tool or argument set Cedar would have denied by sending a permitted request shape that the webhook rewrote into a forbidden one. A second bug narrowed this in practice: r.ContentLength was not refreshed alongside r.Body, so a mutation that shrank the body failed at the reverse proxy and one that grew it was truncated into invalid JSON. The bypass was live for length-preserving rewrites — which is exactly case/format normalization, and a webhook can pad JSON whitespace to hold length constant. That stale Content-Length is also fixed here.

Before

client request  ──► ParsingMiddleware ──► parse cached ──► mutating webhook rewrites body
                                                │                      │
                                          Cedar reads ◄────────────────┘  (pre-mutation)
                                          audit reads                     backend runs post-mutation

After

client request  ──► ParsingMiddleware ──► parse cached ──► mutating webhook rewrites body
                                                                        │
                                                       RepublishParsedMCPRequest (body changed)
                                                                        │
                                          Cedar reads ◄────────────────┘  (post-mutation)
                                          audit reads                     backend runs post-mutation

Migration steps

  1. Check whether you set --webhook-config with a mutating: entry (or MCPWebhookConfig.spec.mutating). If not, stop — no action needed.
  2. Read each mutating webhook's patch and enumerate what it rewrites: the JSON-RPC method, params.name, and/or params.arguments.
  3. Re-check your Cedar policies against the post-mutation shape — resource names (MCP::Tool::"<name>") and every when { context.arg_* } clause. Policies that were passing only because they never saw the rewrite will now deny, and vice versa.
  4. Update SIEM rules, dashboards, and saved queries keyed on audit type or target.name — for mutated requests those values change on upgrade.
  5. Expect two new fail-closed responses replacing what previously reached the backend: 400 if a webhook rewrites a single request into a JSON-RPC batch, and 500 if a webhook emits a body that is not a valid JSON-RPC request.

Gaps this deliberately does not close, all documented rather than fixed:

  • With includeRequestData: true, the recorded request payload is still the pre-mutation body (audit reads r.Body before the webhook), so event type/target name are post-mutation while the payload is not.
  • After a webhook renames a tool, the Mcp-Method/Mcp-Name headers forwarded to the backend still name the original tool. A conformant Modern backend rejects the mismatch, so it fails closed — but a mutating webhook should not rename tools on the Modern path.
  • The tool-call filter and rate limiter run outside ParsingMiddleware and still decide against the request as received, so --tools filtering remains bypassable by a webhook rename. Tracked in #6134.

PR: #6136 — Fixes #6133

Migration guide: Recovered panics are no longer logged

This is an unintended regression, not a design decision. It is called out here because it costs you diagnostics silently, and a one-line fix is expected in a patch release.

Who is affected: any operator who relies on ToolHive's logs to diagnose a recovered HTTP panic — including log-based alerts, log-derived metrics, and support bundles. Everyone running without Sentry configured (the default) is affected most.

What changed: pkg/recovery became a thin shim over toolhive-core/recovery. Core's Middleware recovers panics silently unless a logger is injected via WithLogger, and ToolHive's shim passes only WithPanicHandler. The OTel span error recording and Sentry issue reporting are genuinely preserved — same span status (codes.Error, "panic recovered"), same sanitization, same raw value to Sentry, same ordering — but the slog.Error line and its stack trace are gone, and no other middleware picks them up.

Before (v0.41.0)

time=... level=ERROR msg="Panic recovered: runtime error: index out of range [3] with length 2
Stack trace:
goroutine 42 [running]:
runtime/debug.Stack()
..."

After (v0.42.0)

(no log output — the client receives 500 Internal Server Error and nothing is recorded locally)

Migration steps

  1. If you have alerts or log-based metrics matching Panic recovered, they will stop firing. Do not interpret the silence as "no panics" — re-point them at the 500-response rate or at Sentry until the log line returns.
  2. Configure Sentry if you have not already; ReportPanic still sends the raw panic value, so panics remain visible as Sentry Issues with full context.
  3. Traces are unaffected — the request span still carries RecordError plus an error status, so OTel-based panic detection keeps working.
  4. When the fix lands, the restored line will be structured (msg="panic recovered" with panic, method, path, and stack attributes) rather than the old single formatted string, so write any new log parser against that shape.

PR: #6145

Migration guide: Go API changes

Who is affected: only out-of-tree Go code importing ToolHive packages. No CLI, REST API, or CRD surface changes here, and no in-tree caller is affected.

pkg/telemetry/providers was deleted (#6146)

The package and its /otlp and /prometheus subpackages were removed and consumed from toolhive-core instead. The graduation is verbatim — every non-test file is byte-identical apart from two self-referential import paths — so all 12 options (WithServiceName, WithServiceVersion, WithOTLPEndpoint, WithHeaders, WithInsecure, WithCACertPath, WithTracingEnabled, WithMetricsEnabled, WithSamplingRate, WithEnablePrometheusMetricsPath, WithCustomAttributes, WithExtraSpanProcessors) plus NewCompositeProvider, ProviderOption, and CompositeProvider keep identical names and signatures. Nothing about emitted telemetry changes — resource attributes, service-name defaulting, OTLP exporter/TLS config, and Prometheus exporter registration all behave as before.

Before

import "github.com/stacklok/toolhive/pkg/telemetry/providers"

After

import "github.com/stacklok/toolhive-core/telemetry/providers"

Two optimizerdec constants were removed (#6175)

pkg/vmcp/session/optimizerdec no longer exports CallToolArgToolName or CallToolArgParameters. Both have been part of the published API since v0.15.0. They existed to read the call_tool target out of a raw arguments map, a pattern that is now known-unsafe: encoding/json falls back to case-insensitive field matching, so a map index and a struct decode resolve different key sets.

Before

toolName, _ := args[optimizerdec.CallToolArgToolName].(string)
params, _ := args[optimizerdec.CallToolArgParameters].(map[string]any)

After

// Decode with the same call both dispatch sites use, so key matching cannot diverge.
in, err := schema.Translate[optimizer.CallToolInput](args)

registry.Provider gained three methods (#6135)

ListAvailablePlugins(), GetPlugin(namespace, name), and SearchPlugins(query) were added to the interface. Implementations that embed registry.BaseProvider pick up no-op defaults and need no change; anything satisfying the old method set directly will no longer compile.

Migration: embed registry.BaseProvider in your provider struct, or implement the three methods.

🔄 Deprecations

  • pkg/audit's MCP event constants, LevelAudit, and NewAuditLogger are now transitional aliases for github.com/stacklok/toolhive-core/audit and will be removed once the migration's cleanup wave rewrites imports per subtree — prefer the toolhive-core/audit symbols in new code (#6148)

🆕 New Features

  • Manage plugins for AI coding tools with the new thv ai-plugin command group — build, validate, push, install, list, info, uninstall, plus local build management via builds and builds remove — targeting Claude Code and Codex (#5782)
  • The same plugin surface is available over REST at /api/v1beta/plugins (10 endpoints) with a matching Go HTTP client in pkg/plugins/client, so the CLI, API, and external tooling share one contract (#5782)
  • thv ai-plugin install <name> now resolves a plain name against the configured registry instead of failing with a 404 hint, and new catalog routes let you browse and search plugins in a registry (#6135)
  • Project-scoped skill installs now verify Sigstore signatures before anything is extracted or recorded, recording the signer identity as lock-file provenance: on first use and rejecting unsigned artifacts unless you pass --allow-unsigned (#6129)
  • thv skill sync re-verifies each managed skill's stored Sigstore bundle offline against the lock file's recorded identity, treating a failed re-verification as drift so a CI gate catches signature changes exactly like content changes (#6131)
  • thv skill upgrade refuses to move a skill to an artifact signed by a different identity — or to an unsigned one — reporting signer-change-blocked unless you explicitly rotate trust with --allow-signer-change (#6132)
  • Git-installed skills get full gitsign commit-signature verification, with the chain of trust checked against embedded Fulcio roots and no network access (#6121, #6091)

The skills signing features above are all behind the experimental TOOLHIVE_SKILLS_LOCK_ENABLED gate and apply only to project-scoped installs. With the gate unset, thv skill install behaves exactly as in v0.41.0. Note that git (gitsign) provenance is recorded as provisional: true because the embedded Rekor transparency-log proof is not yet validated — signing time is checked only against the Fulcio certificate's own ~10-minute validity window. OCI provenance is not provisional.

🐛 Bug Fixes

  • Multiple MCP clients can now connect to a single stdio MCP server through ToolHive — the first handshake is cached and replayed instead of every client after the first getting duplicate "initialize" received, which also unblocks vMCP aggregating stdio backends (#6153)
  • A client that retries initialize on a live connection behind the transparent proxy now receives a fresh session instead of a hard failure, because the proxy no longer forwards a session ID on initialize (#6152)
  • vMCP gateways aggregating a dual-era backend such as github-mcp-server v1.6.0 no longer oscillate between the Modern and Legacy revisions and fail roughly half their health checks — a Modern promotion must now win a confirming server/discover probe rather than trusting the negotiated version alone (#6158)
  • A vMCP backend redeployed from a hint-lying Legacy server to a genuinely Modern one now corrects its reported MCP revision within ~5 minutes instead of staying Legacy until the pod restarts (#6185)
  • vMCP now relays backend log and progress notifications from Modern (2026-07-28) backends to the downstream client, opting in through the per-request io.modelcontextprotocol/logLevel _meta key that replaced the removed logging/setLevel RPC (#6140)
  • vMCP clients on the Legacy revision now receive non-reserved backend _meta (trace ids, custom fields) on resources/read results, matching what the Modern path already delivered (#6180)
  • A vMCP pod with the optimizer enabled no longer permanently loses its tool index while continuing to report itself healthy — the in-memory SQLite database is now pinned alive by a dedicated connection, so one cancelled request can't destroy it for the life of the process (#6157)
  • find_tool's tool_keywords input now actually affects results instead of being decoded and dropped, and it drives the lexical BM25 arm while tool_description drives semantic matching (#6124)
  • call_tool now accepts the common LLM malformation where tool_name is nested inside parameters, and a genuinely missing tool_name produces an error that states the expected shape and lists the parameter names received (#6150)
  • On Windows, the discovery directory and server.json under %LOCALAPPDATA% are now protected with an explicit DACL granting only the ToolHive user and SYSTEM, and are ownership-validated before being trusted — POSIX mode bits are advisory on NTFS, so any local account with Modify could previously rewrite the npipe:// discovery URL and redirect the next MCP client (#5951)
  • The authorization middleware now resolves a call_tool target through the same decoder dispatch uses, closing three case-sensitivity divergences that could skip a policy check or drop arguments (#6175)

🧹 Misc

  • pkg/telemetry/providers (~2,900 LOC) is deleted in favour of the verbatim graduation in toolhive-core, with no change to emitted telemetry (#6146)
  • pkg/recovery becomes a thin shim over toolhive-core/recovery, keeping ToolHive's OTel and Sentry wiring through a panic-handler hook (#6145)
  • MCP histogram buckets are sourced from toolhive-core's semconv preset instead of a local literal; the boundaries are unchanged (#6144)
  • Audit event constants, LevelAudit, and NewAuditLogger become aliases over toolhive-core/audit with byte-identical values, so the audit wire format is untouched (#6148)
  • Pinned a regression test for vMCP elicitation failing fast when the client advertised the capability but holds no standalone SSE stream, and documented both delivery constraints (#6182)
  • Fixed a port-selection TOCTOU race that flaked e2e tests under sharded CI by having the OIDC and LLM gateway mocks hold their own listener from construction (#6142)
  • Pinned the ida-pro-mcp e2e image by digest after an upstream rebuild pulled in the breaking mcp Python SDK 2.0.0, and added test/e2e/images/** to the lifecycle suite's trigger filter so an image change can no longer skip the tests that consume it (#6159)
  • Pinned the mcp-server-time e2e image by digest for the same upstream breakage, unblocking the proxy suites (#6160)
  • Fixed the operator integration suites' timeout waiting for process kube-apiserver to stop flake — 32 of the job's last 51 failures — by awaiting manager shutdown before tearing down envtest (#6179)

📦 Dependencies

Module Version
github.com/stacklok/toolhive-core v0.0.35 → v0.0.38
github.com/stacklok/toolhive-catalog v0.20260804.0
github.com/tailscale/hujson b80ff77
coverallsapp/github-action 8d6379e
github/codeql-action f205ea1
anthropics/claude-code-action v1.0.183

toolhive-core was bumped across #6144, #6146, and #6180 rather than by a dependency PR; v0.0.38 also carries transitive bumps to aws-sdk-go-v2, go-containerregistry, moby/client, prometheus, and otel.

👋 Welcome to our newest contributor: @Tanguille 🎉

Full commit log

What's Changed

New Contributors

Full Changelog: v0.41.0...v0.42.0

🔗 Full changelog: v0.41.0...v0.42.0

@stantheman0128

Copy link
Copy Markdown
Contributor

Nice to see #5951 ship in v0.42.0. Thanks again to @samuv for the thorough security review rounds on the Windows discovery DACL work — glad the LOCALAPPDATA hardening made this release.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release size/XS Extra small PR: < 100 lines changed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants