Skip to content

test(chunk-grids): sample the full rectilinear chunk grid declaration space - #4377

Merged
d-v-b merged 100 commits into
zarr-developers:mainfrom
d-v-b:test/rectilinear-declaration-space
Oct 2, 2026
Merged

d-v-b merged 100 commits into
zarr-developers:mainfrom
d-v-b:test/rectilinear-declaration-space

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

🤖 AI text below 🤖

The hypothesis strategies in zarr.testing.strategies now sample the full space of rectilinear chunk grid declarations: arrays() passes a rectilinear grid to create_array either as metadata or as a chunks= specification mixing bare-int steps and explicit edge lists in any arrangement, and the new rectilinear_chunk_shape_declarations and rectilinear_chunk_grids cover stored metadata, including run-length encoded edge lists in canonical and arbitrary groupings and edges overhanging the array extent. rectilinear_chunks and chunk_grids also draw zero-length dimensions.

Follow-up to #4374. Tests and zarr.testing strategies only; additive.

Why

The bug in #4374 shipped in 3.2.x because the rectilinear hypothesis strategy returned one list of edges per dimension, so no property test ever saw a bare-int dimension, a grid mixing bare ints and edge lists, a run-length encoded declaration, or edges overhanging the extent. A classifier that only looked at the first element was invisible to it.

Changes

The strategies now cover the two spaces the spec has: the chunks= input syntax, and the stored chunk_shapes metadata.

  • chunks= space. arrays() passes a rectilinear grid to create_array either as the metadata object or as a chunks= specification in which each dimension is a bare-int step (which may exceed the extent) or a flat edge list summing to the extent, in any arrangement, with at least one list. Run-length encoding is not part of this space: it is a metadata form, not an input form. arrays() asserts that the stored grid equals the declared one (bare ints stay bare ints, edge lists keep their edges).
  • Stored space. New rectilinear_chunk_shape_declarations(shape=...) returns a stored chunk_shapes value and the grid it must parse to: bare-int steps, edge lists written in full or run-length encoded (canonical, or any grouping, including split runs and [size, 1] pairs), and edges overhanging the extent (a trailing edge, or a last edge past the end). New rectilinear_chunk_grids(shape=...) draws RectilinearChunkGridMetadata from the same space, and chunk_grids (so array_metadata) uses it.
  • Zero-length axes are drawn in both spaces. On an axis of extent 0 an edge list is any non-empty list of positive edges: the chunks the axis grows into, the state zarr 3.2.x wrote, resize(0) produces and fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334 lets chunks= create. chunk_grids used to fall back to a regular grid for any empty axis. rectilinear_dim_edges(extent=...) is the shared per-axis draw.
  • Public API kept. rectilinear_chunks keeps its 3.4.0 form (one edge list per dimension, [] for a 0-d shape) and now also draws zero-length dimensions; chunks_param_from_rectilinear is kept. Their changes (bare-int steps in rectilinear_chunks, removing chunks_param_from_rectilinear) belong to the 3.5.0 follow-up. The per-dimension chunk limit and the overall chunk budget are declared once.
  • tests/test_unified_chunk_grid.py and fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334's state machine (tests/test_array_stateful.py) drop their private copies of the strategy and draw from the shared one.

New properties (tests/test_properties.py)

  • every stored declaration parses to its expanded edges and re-serializes to an equal grid, also inside a whole ArrayV3Metadata document;
  • every chunks= specification lands in zarr.json as a "rectilinear" grid whose chunk_shapes equal the specification (compared against an independent run-length expansion of the stored JSON, not against the class under test);
  • rectilinear_chunks declares one edge list per dimension;
  • a rectilinear array with a zero-length axis round-trips data through append and resize.

Reach

Measured with --hypothesis-show-statistics (default profile, 300 examples, share of examples):

Test Zero-extent edge list Step larger than extent Overhang Arbitrary RLE grouping
stored-declaration property 5.0% 22.4% 27.6% 25.2%
chunks= property 11.0% 10.2%
zero-length round trip 88.1% 27.8%
test_array_roundtrip 1.6% 1.6% 3.0%
test_basic_indexing 8.2% 1.2%

Every one of these was 0% before. No assume() was added. Block indexing has no zero-length rectilinear shapes, since it needs at least one chunk per axis.

Stack and release notes

🤖 Generated with Claude Code

…, clamps and metadata

Invariant: a chunk edge length is always >= 1; a dimension's extent may be 0,
in which case the dimension has zero chunks (ceildiv(0, size) == 0).

Zero-length-axis bugs have recurred since 2017 (#150, #241, #303, zarr-developers#972,
zarr-developers#1977, zarr-developers#2434, zarr-developers#3711, zarr-developers#4305, zarr-developers#4307, zarr-developers#4328) because the layers disagreed on
this invariant and every span-derived chunk spelling clamped on its own:

- The metadata layer (common.py, metadata/v3.py) required chunk edges >= 1,
  but the in-memory FixedDimension allowed size == 0 with four special-case
  branches left over from zarr-developers#2434, so normalization could build a grid the
  metadata constructor then rejected. FixedDimension now rejects size < 1
  and the four `if self.size == 0` branches are gone. VaryingDimension
  already required edges > 0 and is unchanged.
- `chunks=-1`, `chunks=False`, `chunks="auto"` (_guess_regular_chunks, both
  the typesize == 0 early return and the np.maximum line) and `shards="auto"`
  each derived "one chunk covering the axis" independently. They now all go
  through one helper, `_full_span_chunk_size(span) = max(span, 1)`, which is
  the single definition of that phrase for a possibly zero-length axis.
- Zarr format 2 metadata had no chunk >= 1 check, so a legacy `chunks: [0]`
  document opened fine and read uninitialised memory after a resize. It now
  raises a clear ValueError at parse time, matching the format 3 grid.
- Rectilinear grids had no creation-time spelling for a zero-length axis:
  normalize_chunks_1d required sum(edges) == span, which no list of positive
  edges can satisfy for span 0, even though the same state is reachable via
  resize((0,)) and round-trips through reopen. For span == 0 any non-empty
  list of positive edges is now accepted verbatim, producing the same
  VaryingDimension(edges, extent=0) that resize produces; the strict sum
  check is kept for span > 0.

Tests: the per-spelling regression test from zarr-developers#4328 is replaced by one matrix
over {-1, False, "auto", 1, (1,...), [[2, 2]]} x {(0,), (0, 4), (4, 0),
(0, 0), ()} x {v2, v3} x {no shards, shards="auto" with and without a byte
budget, explicit shards}, with separate small tests for each error case.
Tests that constructed FixedDimension(size=0) now assert it raises, and a
zero-extent test covers the behaviour the old special cases were guarding.

Assisted-by: ClaudeCode:claude-fable-5-1
zarr-python 2.18.7 writes `chunks: [0]` for `zarr.zeros((0,), chunks=False)`
and for `chunks=(0,)`, so stores with that document exist. Rejecting them
at open would turn a previously-readable array into an error; leaving the
0 in place read uninitialised memory after a resize. Normalize the edge to
1 with a ZarrUserWarning instead — the same grid every other "one chunk
spans the axis" spelling produces — and keep rejecting a zero edge on an
axis that has data.

Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
Measured against zarr 2.18.7: `zeros((0,), chunks=False)`, `chunks=-1` and
`chunks=(0,)` all write `chunks: [0]`, after which nchunks, read, write,
append, resize and reopen-then-read every raise ZeroDivisionError. There
was never a working behaviour to preserve; normalizing the edge to 1 makes
such arrays usable for the first time. Say so in the comment and fragment
instead of claiming the stores were previously readable.

Assisted-by: ClaudeCode:claude-fable-5-1
The rectilinear hypothesis strategy returned one list of edges per
dimension, so no property test ever saw a bare-int dimension, a grid
mixing bare ints and edge lists, a run-length encoded declaration, or
edges overhanging the extent. The first-element-only classifier that
zarr 3.2.x shipped (zarr-developers#4374) was invisible to it.

The strategies now cover two spaces. `rectilinear_chunks` samples the
`chunks=` syntax: bare ints and flat edge lists in any arrangement, at
least one list so the grid is rectilinear. `rectilinear_chunk_shape_
declarations` samples the stored metadata: bare-int steps (including
larger than the extent), edge lists written in full or run-length
encoded in canonical or arbitrary grouping, and overhanging edges. Each
draw comes with the chunk_shapes it must parse to. `chunk_grids` and so
`array_metadata` draw from the stored space; `arrays` passes grids the
list syntax cannot express as the metadata object, and asserts the
stored grid equals the declared one.

Two property tests pin the properties that would have caught zarr-developers#4374:
every stored declaration parses to its expanded edges and re-serializes
to an equivalent grid, and every `chunks=` specification is stored in
zarr.json as a "rectilinear" grid equal to the specification.

test_unified_chunk_grid.py used a private copy of the old strategy; it
now draws from the shared one.

Assisted-by: ClaudeCode:claude-fable-5-1
d-v-b added a commit to d-v-b/zarr-python that referenced this pull request Sep 18, 2026
Assisted-by: ClaudeCode:claude-fable-5-1
@github-actions github-actions Bot added the needs release notes Automatically applied to PRs which haven't added release notes label Sep 18, 2026
Assisted-by: ClaudeCode:claude-fable-5-1
@github-actions github-actions Bot removed the needs release notes Automatically applied to PRs which haven't added release notes label Sep 18, 2026
@read-the-docs-community

read-the-docs-community Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

@read-the-docs-community

read-the-docs-community Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Documentation build overview

📚 zarr-indexing | 🛠️ Build #34756226 | 📁 Comparing 84b19cc against latest (1187a43)

  🔍 Preview build  

No files changed.

@codecov

codecov Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.66%. Comparing base (89f058c) to head (2976502).

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #4377      +/-   ##
==========================================
+ Coverage   94.63%   94.66%   +0.02%     
==========================================
  Files          94       94              
  Lines       13487    13544      +57     
==========================================
+ Hits        12763    12821      +58     
+ Misses        724      723       -1     
Files with missing lines Coverage Δ
src/zarr/testing/strategies.py 97.03% <100.00%> (+0.40%) ⬆️

... and 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

… format 3

The compatibility policy for a stored chunk size of 0 on a zero-length
axis covered Zarr format 2 only, so arrays written by zarr-python 3.0 and
3.1 with `chunk_shape: [0]` — or `[false]`, which 3.0 wrote for
`chunks=False` — still could not be opened at all.

`ArrayV3Metadata` now applies the same policy to a regular chunk grid: a
stored chunk size of 0 on a zero-length axis is read as 1 with a
`ZarrUserWarning`, and a zero chunk size on a positive-length axis is
left for the chunk grid parser to reject. It runs in `__init__` rather
than in the grid parser because the policy needs the array shape, which
chunk grid metadata does not carry.

Both warnings now say how to store a corrected chunk size — open the
array writable and call `array.update_attributes({})`, which rewrites the
whole document from the parsed metadata — from one shared constant. The
Zarr format 2 warning also named only zarr-python 2.x; measured against
real installs, every 3.x release before 3.4 wrote a zero chunk size for
an empty array too (3.3.0 for `chunks=-1` and `chunks=False`).

Tested against stores written by zarr 3.0.10, 3.1.6 and 3.3.0: they open,
append without losing data, and re-save to a chunk size that reopens
without a warning.

Assisted-by: ClaudeCode:claude-opus-5
d-v-b and others added 11 commits September 19, 2026 18:05
… in one routine

The legacy zero-chunk policy was written out twice: inline in
`ArrayV2Metadata.__init__`, and again in a Zarr format 3 helper. The V2
copy also zipped with `strict=False` and re-appended any trailing chunk
entries only so that a separate length check, `parse_metadata`, could
report a dimensionality mismatch after construction.

`parse_stored_chunk_shape` in `zarr.core.metadata.common` is now the one
place a stored chunk shape is checked against its array's shape, for both
formats: one entry per axis, every integer chunk size at least 1, and a
size of 0 (or JSON `false`) on a zero-length axis read as 1 with a
warning that names the writer and how to re-save. Non-integer entries,
such as edge lists, pass through for the caller's own parser.

`ArrayV2Metadata.__init__` calls it directly and `parse_metadata` is
gone. The Zarr format 3 adapter only locates a regular grid's
`chunk_shape` in the stored document and hands it over; it still runs in
`ArrayV3Metadata.__init__` because chunk grid metadata has no array shape.
The Zarr format 3 import changes that existed only for the old helper are
reverted.

Tests for the policy now target the routine: one table of valid and
legacy inputs, and one test per rejection (dimension mismatch, zero on a
non-empty axis, negative). They replace metadata-level tests in
test_v2.py and test_v3.py that only re-tested the same rules; the
end-to-end tests still cover both formats' wiring against stored arrays.

Assisted-by: ClaudeCode:claude-opus-5
…nk grids

`parse_stored_chunk_shape` passed non-integer entries through "for the
caller's own parser", which made a regular-grid policy look like a
general chunk shape routine and let it decide what a 0-length chunk means
for grids it does not own. A rectilinear grid, or any other grid, is free
to define its own semantics for 0-length chunks.

It is now `parse_stored_regular_chunk_shape`, typed `Sequence[int]`, with
no pass-through, and its docstring says it applies to Zarr format 2
`chunks` and Zarr format 3 `regular` grids only. The Zarr format 3 caller
hands it a chunk shape only when the grid is named `regular` and every
entry is an integer (`_is_regular_chunk_shape`); anything else is not a
regular chunk shape and goes to the chunk grid parser untouched.

Assisted-by: ClaudeCode:claude-opus-5
- Document the `rectilinear_chunks` return type change in the changelog
  fragment, since `zarr.testing.strategies` is public.
- Make the `RectilinearDimDeclaration` alias private.
- Encode canonical RLE in the strategy instead of calling `compress_rle`,
  so the generator does not depend on the code under test.
- Assert `rectilinear_chunks` gets a non-empty shape instead of returning
  a non-rectilinear `[]`.

Assisted-by: ClaudeCode:claude-opus-5-5
# Conflicts:
#	tests/test_properties.py
With zarr-developers#4334 a rectilinear dimension of extent 0 takes any non-empty list of
positive edges at creation, the same state 3.2.x wrote and resize(0)
produces. The strategies asserted extent > 0 and `chunk_grids` fell back to
a regular grid for any empty axis, so no property saw that state.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…panning chunk

A stored chunk size of 0 was tolerated only on a zero-length axis and
rejected otherwise, because the metadata supposedly could not say how the
stored chunks were laid out. But a chunk size of 0 gives a grid of zero
chunks, so no release could store a chunk under it, and the writers of
that metadata let the axis grow: 3.4.0 appends to a Zarr format 2 array
created empty by 3.3.0 (shape grows, no chunk written), and 3.1.6 records
a Zarr format 3 resize and the resize half of a failed append. Measured
with real installs. Those arrays open in 3.4.0, attributes included; the
rejection would have made them unopenable.

`parse_stored_regular_chunk_shape` now reads a stored 0 (or JSON `false`)
on any axis as one chunk spanning it, `max(extent, 1)`, which is what the
`-1`/`False` spec that wrote it meant. On a grown axis the warning also
says that data written to it was not saved. Negative sizes are still
rejected.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The only state machine that touched arrays compared zarr on one store with
zarr on a MemoryStore, so a chunk grid bug showed up identically on both
sides; it had no append rule, covered Zarr format 3 only, and kept every
empty axis at 0 when resizing, which is where the zero-length bugs live.

`ArrayLifecycle` checks one array against a NumPy model across both
formats, every chunk spelling (-1, False, "auto", ints, sharded,
rectilinear) and the stored chunk size of 0 that releases before 3.4
wrote, including on an axis those releases grew. Rules append, resize
(growing and shrinking to and from 0), write and re-save the metadata;
the invariant reopens the array and compares shape, values and whether
the legacy warning is due. Deliberately breaking the grown-axis policy,
the legacy warning, or append on an empty axis each fails it.

`resize` keeps partly retained chunks whole, so cells cut off by a shrink
can come back with their old values when the axis grows (as in 2.x); the
model marks such cells unknown until written instead of encoding that.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The warning for a stored chunk size of 0 said which zarr-python releases
wrote it. The check runs on every metadata construction, so metadata built
in code, such as VirtualiZarr's kerchunk writer passing an empty array's
shape as `chunks`, was told it came from zarr-python 2.x. The warning now
says what the chunk size is read as and, on an axis of positive length,
that the axis holds only the fill value. `legacy_writers` is gone, and
the docstrings no longer narrate release history; the changelog keeps it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
d-v-b and others added 4 commits September 26, 2026 20:23
…tadata

Setting attributes of an array read from a document that had to be
upgraded stored the upgraded document with the new attributes, but left
the metadata marked with the document it was read from. A consolidated
group handle shares that metadata, so its next write stored the old
document in the consolidated metadata and the new attributes were lost
from it (zarr 3.4.0 kept them).

Every write of an array's own documents (creating, resizing, setting
attributes) now goes through `AsyncArray._save_metadata`, which clears
the mark afterwards, as the first chunk write does after storing the
upgrade. Both clear it through `_stored_document_replaced`, the one
place that declares it: the first chunk write clears it also when it
stores nothing (the store already holds a valid document, or none), so
it cannot be folded into the save.

The deep copy in `mark_upgraded` stays: the metadata's attributes share
objects with the document it was read from, which the caller also
holds. A test pins it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ment mark

Setting attributes on or resizing an array read from a document that needed
an upgrade clears the `_stored_document` mark only after the store accepts the
new metadata. A test now injects a store failure for the array's own document
and checks the mark and the stored bytes survive, for Zarr formats 2 and 3.
The `_save_metadata` docstring now says only what the method does.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@d-v-b d-v-b changed the title test: sample the full rectilinear chunk grid declaration space test(chunk-grids): sample the full rectilinear chunk grid declaration space Sep 29, 2026
A stored chunk size of 0 or `false` was read as one chunk spanning the
axis at open time. On an axis that grew after the size was written, that
made the chunk size depend on when the array was opened, and it could be
very large: `chunks: [0, 100, 100]` on shape (10000, 100, 100) read as one
800 MB chunk. It is now read as the smallest valid chunk edge length (1,
or the inner chunk size for a shard), which is what `chunks=-1` gives on
a zero-length axis. A resize by software that kept the stored 0 no longer
changes the layout a stale handle expects.

Assisted-by: ClaudeCode:claude-opus-5-5
…SON input as is

A stored `true` chunk size or an integral float edge is read as the value
zarr already read it as, so chunks written under the upgraded metadata are
where every reader looks. Marking such documents for re-saving turned a
plain chunk write into a metadata write: in a ZipStore that adds a second
`zarr.json` entry (a `UserWarning`, an error under `-W error`), and it
opened a race with concurrent metadata writes that 3.4.0 did not have. Each
upgrade now reports whether it moves chunks, and only those mark the
metadata.

The upgrades also detected changes by comparing JSON encodings, so
`ArrayV3Metadata.from_dict` given metadata built in code (a codec instance,
a NumPy integer in a sharding codec's `chunk_shape`) raised `TypeError`.
Changes are now tracked where they are made.

Assisted-by: ClaudeCode:claude-opus-5-5
@d-v-b d-v-b added this to the 3.4.1 milestone Sep 29, 2026
d-v-b and others added 11 commits September 30, 2026 10:44
…was 0

The warning for a stored chunk size of 0 on an axis of positive length now says that
the array holds only its fill value, so recreating it with the wanted chunk shape loses
nothing, and gives the re-save as the way to keep it instead. A re-save freezes the
smallest chunk size into an array that holds no data yet. Each reading now carries its
own advice, so mark_upgraded appends no shared hint.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ored chunk size of 0

The warning names the call, zarr.from_array(array.store, name=array.path, data=array,
chunks=..., overwrite=True, write_data=False), which keeps the data type, fill value,
attributes and codecs, and a test pins the recipe on both formats.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Resolves conflicts with zarr-developers#4410, which replaced _prepare_overwrite with
save_new_metadata: keep main's helper and this branch's imports and chunk
normalization.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…airs, not upgrades

The module is zarr.core.metadata.repair: repair_array_document, mark_repaired,
ARRAY_REPAIRS and the Repair type, AsyncArray._store_repaired_document, and the test
module test_repair.py. A reading that turns an invalid stored document into a valid one
fixes it; it does not move it to a newer format.

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
main now holds zarr-developers#4334 as a squash commit whose tree equals the zarr-developers#4334 head this branch
already contains, so every file it touches keeps this branch's version; the merge
brings in only the dependency bump (zarr-developers#4461).

Assisted-by: ClaudeCode:claude-fable-5-1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@d-v-b

d-v-b commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

this only hits our test suite so i'm self-merging

@d-v-b
d-v-b marked this pull request as ready for review October 2, 2026 07:37
@d-v-b
d-v-b merged commit 7db5831 into zarr-developers:main Oct 2, 2026
39 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant