Repository navigation
Remove terraform dependency from invariant tests - #6876
Merged
Merged
Conversation
invariant/auto-migrate and invariant/migrate seeded their migration by deploying on the terraform engine. So the engine can be removed, they now replay a per-config terraform-state fixture (captured once from a real terraform run) via replay_tfstate.py, like the migrate/ tests. They are local-only (Cloud=false): the fixtures are captured against the fake server, so they reproduce a terraform-shaped state only there. Generalize the capture/replay tooling for the fuzzer's resource diversity: - capture_tfstate.py keeps workspace/mkdirs (a dashboard's parent_path needs it) and the recorded query params (q), which carry required ids for some creates (postgres ?project_id=...); adds --out for per-config fixtures + target auto-detection (the invariant configs declare no target). - replay_tfstate.py reattaches those query params, matches clusters by cluster_name / remaps cluster_id, and skips the id-remap when a resource's id is the identifying name it already sent (registered models, serving endpoints); otherwise it mapped the name to a separate backend uuid and corrupted the id. Excluded (fixtures captured, awaiting follow-up): model_with_permissions and model_serving_endpoint; their permissions reference a backend-minted secondary id (mlflow registered_model_id / serving endpoint id) that the fake server's create response omits, so replay cannot remap it. model.yml keeps migrate coverage for registered models. Co-authored-by: Isaac <no-reply@databricks.com>
Collaborator
Integration test reportCommit: 7364eb9
Top 6 slowest tests (at least 2 minutes):
|
… perms migrate
The invariant fuzzer had excluded model_with_permissions + model_serving_endpoint:
their permissions target a backend-minted SECONDARY id (the mlflow
registered_model_id / the serving-endpoint id), distinct from the resource's
terraform id (its name), which replay was not remapping - so the permission
landed on the stale recorded id and showed as drift.
replay_tfstate.py now, for a create whose primary id is the client-provided
name, maps the recorded secondary id to the freshly minted one: it either rides
the create response as `id` (serving endpoints), or comes from the databricks
registered-models GET that terraform issues (mlflow models - the create response
returns only the OSS model, matching the real API, and capture drops GETs). The
fake server already mints these ids, so it stays faithful - no server change.
Re-enables both configs in invariant/{migrate,auto-migrate}.
Co-authored-by: Isaac <no-reply@databricks.com>
…two configs Removing the model_with_permissions / model_serving_endpoint EnvMatrixExclude entries changes the resolved matrix, so out.test.toml (which the harness rewrites at runtime) has to be regenerated - otherwise the post-test "no files changed" check fails. Co-authored-by: Isaac <no-reply@databricks.com>
denik
marked this pull request as ready for review
September 30, 2026 08:43
denik
enabled auto-merge
September 30, 2026 08:43
The prior comment said the migrate/auto-migrate invariant tests "cannot run against a real workspace", which overstates it. Explain the actual reasons: the per-config terraform-state fixture is captured against the local test server, and replay_tfstate.py authenticates its raw API calls with $DATABRICKS_TOKEN, which is empty under cloud OAuth. A cloud path is possible but needs SDK-chain auth in replay (via `$CLI api`, once the cmd/api int64-precision bug is fixed) plus fixtures captured against a real workspace. Co-authored-by: Isaac <no-reply@databricks.com>
Reword the comment: these run local-only permanently, not "for now". tf->direct migration is a one-time transition over a closed set of resource types -- terraform gains no new converters, and resources created on the direct engine are never migrated -- so replaying the captured fixtures against the local test server is complete coverage and a real-workspace run would add nothing. Co-authored-by: Isaac <no-reply@databricks.com>
shreyas-goenka
approved these changes
Sep 30, 2026
janniklasrose
approved these changes
Sep 30, 2026
Collaborator
Integration test reportCommit: 5b4efa9
524 interesting tests: 524 FAIL
Top 50 slowest tests (at least 2 minutes):
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
invariant/auto-migrateandinvariant/migrateseeded their migration by deploying on the terraform engine. So the engine can be removed, they now replay a per-config terraform-state fixture (captured once from a real terraform run) viareplay_tfstate.py, like themigrate/tests. They run local-only, which is complete coverage rather than a gap: tf->direct migration is a one-time transition over a closed set of resource types (new resources are created on the direct engine and never migrated), so replaying the captured fixtures against the fake server covers it fully.Generalizes the capture/replay tooling for the fuzzer's resource diversity:
workspace/mkdirs(a dashboard'sparent_path) and the recorded query params (postgres?project_id=...) in the captured fixtures;cluster_name/ remapcluster_id;registered_model_id, serving endpoint id) -- for mlflow, fetched from the databricks registered-models GET the way terraform does. This letsmodel_with_permissionsandmodel_serving_endpointmigrate.This pull request and its description were written by Isaac.
Related: #6866