Skip to content

fix(chunk-grids): read mixed regular/rectilinear chunk grids written by zarr 3.2.x - #4375

Draft
d-v-b wants to merge 96 commits into
zarr-developers:mainfrom
d-v-b:fix/mixed-regular-chunk-grid-4374
Draft

d-v-b wants to merge 96 commits into
zarr-developers:mainfrom
d-v-b:fix/mixed-regular-chunk-grid-4374

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

🤖 AI text below 🤖

Arrays whose stored regular chunk grid mixes chunk sizes with lists of chunk edge lengths, such as "chunk_shape": [2, [5, 10, 5]] (written by zarr 3.2.0 and 3.2.1 for chunks=(2, (5, 10, 5)), and as [2, [5.0, 10.0, 5.0]] for float edges), can be read again, without enabling array.rectilinear_chunks. The grid is read as the rectilinear chunk grid it describes, with a ZarrUserWarning; re-saving the array's metadata (or writing chunks to it) stores that rectilinear grid, which requires the flag. A group's consolidated metadata keeps such an array's metadata as it was stored. A regular chunk grid given edge lists is now rejected, so this metadata is no longer written.

The array.rectilinear_chunks flag now gates reading and storing array metadata that declares a rectilinear chunk grid, including in a group's consolidated metadata, instead of constructing RectilinearChunkGridMetadata. Array.resize, deleting a group member, and create_hierarchy(..., overwrite=True) now encode the metadata they will store before deleting anything, so metadata that cannot be stored (for example, a rectilinear chunk grid with the flag off) fails with the store untouched. That error names the array when it is a member of a group's consolidated metadata.

Closes #4374.

Problem

With array.rectilinear_chunks enabled, zarr 3.2.0 and 3.2.1 classified a mixed chunk specification such as chunks=(2, (5, 10, 5)) as regular (the classifier looked only at the first element) and stored

{"name": "regular", "configuration": {"chunk_shape": [2, [5, 10, 5]]}}

while laying the chunks out as a rectilinear grid. 3.2.x also wrote [2, [5.0, 10.0, 5.0]] for float edges and [true, [5, 10, 5]] for a True size. The classifier was fixed in #4218, but 3.4.0 still accepts such a grid in RegularChunkGridMetadata, writes it, and then fails with an unrelated TypeError: Expected an iterable of integers when building the codec pipeline, so these arrays cannot be read.

Changes

Read the 3.2.x mixed grid, without the flag

The reading is one more upgrade in zarr.core.metadata.upgrades, the module #4334 introduces as the only lenient reader of stored documents. A stored regular grid whose chunk_shape mixes chunk sizes with flat lists of edges is read as the rectilinear grid it describes; its sizes and edges go through the same readings as any other stored chunk size or rectilinear edge (true → 1, 4.0 → 4). A chunk_shape made only of lists, or with nested lists, was never written by any release and is left for the constructors to reject.

This is a compatibility read, so it does not require array.rectilinear_chunks (3.2.x wrote these with the flag set, but the user reading them now may not have it). It warns, because re-saving needs the flag:

Array 'file:///data/x.zarr': The stored chunk grid is named 'regular', but its chunk shape lists chunk edge lengths in dimensions [1], which only a rectilinear chunk grid can declare. It is read as that rectilinear chunk grid. Re-saving the metadata stores that rectilinear chunk grid, so each step that follows requires zarr.config.set({'array.rectilinear_chunks': True}). To store valid metadata, ...

RegularChunkGridMetadata now rejects an edge list with a TypeError naming the dimension, so this metadata can no longer be written.

The flag gates store boundaries, not a class

In 3.4.0 the flag was checked in RectilinearChunkGridMetadata.__post_init__. It now gates the two places a rectilinear chunk grid crosses the store boundary:

  • Storing: check_storable runs in ArrayV3Metadata.to_buffer_dict and, for every array in a group's consolidated metadata, in GroupMetadata.to_buffer_dict. Every serialization of array metadata for a store passes through one of them.
  • Reading: ArrayV3Metadata.from_dict checks a document whose grid is named rectilinear before any upgrade runs (as on main). The 3.2.x mixed grid is exempt, since its document names a regular grid.

Both raise RectilinearChunksDisabledError, a new subclass of ValueError (so existing except ValueError handlers still catch it), with a note naming the array and saying that nothing was stored or read. A consolidated member that fails the gate is named by its path in the consolidated metadata. Consolidated metadata keeps a mixed-grid member in the form it was stored (#4334's rule for upgraded members), so writing group metadata with the flag off does not fail because of it.

Encode before destroy

Any operation that deletes store content now encodes every document it will write first, through one helper (encode_documents), so metadata that cannot be stored (for example, a rectilinear grid with the flag off) fails with the store untouched:

  • Array.resize encodes the new metadata before deleting chunks outside the new shape. A resize on a store that cannot delete (ZipStore) still fails before storing new metadata, as in 3.4.0.
  • del group[name] on a group with consolidated metadata encodes the group metadata without the member before deleting it. After the deletion it pops the member in place (so every handle sharing that consolidated metadata sees it, as in 3.4.0) and stores the encoding of the metadata as it then is, so concurrent deletions each store the deletions made before them.
  • create_hierarchy(..., overwrite=True) builds and encodes every node before deleting anything.

What users see

Situation 3.4.0 This PR
Open a 3.2.x mixed-grid array (flag off or on) fails later with TypeError: Expected an iterable of integers opens as a rectilinear array, with the warning above
Open a consolidated group with such a member opens opens; the member warns when read
Write chunks to, resize, or update_attributes that array with the flag off fails RectilinearChunksDisabledError naming the array; store untouched
Same with the flag on fails stores the rectilinear grid, then proceeds; the array reopens without a warning
RegularChunkGridMetadata(chunk_shape=(2, (5, 5))) accepted; wrote an unreadable document TypeError: Dimension 1: chunk edge length must be an int, got (5, 5)
RectilinearChunkGridMetadata(...) with the flag off ValueError built; storing or reading it is gated instead

Evidence

  • Stores written by real zarr==3.2.0 and zarr==3.2.1 installs (a mixed grid, a mixed grid resized to 0 along its regular axis, a mixed grid with an empty rectilinear axis, a rectilinear grid on an empty axis) open, take an append without losing data, and re-save with the flag to metadata that reopens without a warning. A test fixture is copied verbatim from a 3.2.1-written document.
  • With the flag off, a write or resize of a mixed-grid array raises and leaves every key in the store byte-identical.
  • The 3.4.0 behaviour table and check_patch.py (OK, 0 undocumented changes, 0 warnings-only changes), the byte-identical write matrix, and the xarray and VirtualiZarr runs described in fix(chunk-grids): require chunk sizes of at least 1 and read the 0, false and true sizes older releases stored #4334 cover this branch.

Tests

  • tests/test_metadata/test_upgrades.py: one table for the mixed-grid reading (where the edge lists sit, true and float sizes, a long edge list), one test per rejection (a grid of only edge lists, run-length pairs, a non-integer edge, an edge below 1, edges that do not sum to the extent), a round trip by re-save and by write, the store left untouched without the flag (resize, write, create_hierarchy(overwrite=True)), consolidating a group with the mixed member at mixed and sub/mixed, and group writes (delete, attributes) storing the member as stored with the flag off and on. The fixture is a document copied verbatim from a store zarr 3.2.1 wrote.
  • tests/test_group.py: del group[name] on consolidated groups of both formats (reopened from the store), an aliased subgroup handle seeing the deletion, and concurrent deletions storing the member list the handle holds.
  • tests/test_store/test_zip.py: a ZipStore resize that needs deletes leaves the array unchanged, with no duplicate entries.
  • tests/test_unified_chunk_grid.py, tests/test_metadata/test_v3.py: a regular grid rejects edge lists; the metadata classes are not gated; a stored rectilinear document read without the flag names the array.

Known limitations

  • Concurrent del group[name] calls on one consolidated group can still store a stale member list, about as often as 3.4.0 does.
  • create_hierarchy given user-built metadata that already holds a 3.2.x mixed grid can store it as given.
Stack and release notes

🤖 Generated with Claude Code

d-v-b added 10 commits September 9, 2026 20:28
…, clamps and metadata

Invariant: a chunk edge length is always >= 1; a dimension's extent may be 0,
in which case the dimension has zero chunks (ceildiv(0, size) == 0).

Zero-length-axis bugs have recurred since 2017 (#150, #241, #303, zarr-developers#972,
zarr-developers#1977, zarr-developers#2434, zarr-developers#3711, zarr-developers#4305, zarr-developers#4307, zarr-developers#4328) because the layers disagreed on
this invariant and every span-derived chunk spelling clamped on its own:

- The metadata layer (common.py, metadata/v3.py) required chunk edges >= 1,
  but the in-memory FixedDimension allowed size == 0 with four special-case
  branches left over from zarr-developers#2434, so normalization could build a grid the
  metadata constructor then rejected. FixedDimension now rejects size < 1
  and the four `if self.size == 0` branches are gone. VaryingDimension
  already required edges > 0 and is unchanged.
- `chunks=-1`, `chunks=False`, `chunks="auto"` (_guess_regular_chunks, both
  the typesize == 0 early return and the np.maximum line) and `shards="auto"`
  each derived "one chunk covering the axis" independently. They now all go
  through one helper, `_full_span_chunk_size(span) = max(span, 1)`, which is
  the single definition of that phrase for a possibly zero-length axis.
- Zarr format 2 metadata had no chunk >= 1 check, so a legacy `chunks: [0]`
  document opened fine and read uninitialised memory after a resize. It now
  raises a clear ValueError at parse time, matching the format 3 grid.
- Rectilinear grids had no creation-time spelling for a zero-length axis:
  normalize_chunks_1d required sum(edges) == span, which no list of positive
  edges can satisfy for span 0, even though the same state is reachable via
  resize((0,)) and round-trips through reopen. For span == 0 any non-empty
  list of positive edges is now accepted verbatim, producing the same
  VaryingDimension(edges, extent=0) that resize produces; the strict sum
  check is kept for span > 0.

Tests: the per-spelling regression test from zarr-developers#4328 is replaced by one matrix
over {-1, False, "auto", 1, (1,...), [[2, 2]]} x {(0,), (0, 4), (4, 0),
(0, 0), ()} x {v2, v3} x {no shards, shards="auto" with and without a byte
budget, explicit shards}, with separate small tests for each error case.
Tests that constructed FixedDimension(size=0) now assert it raises, and a
zero-extent test covers the behaviour the old special cases were guarding.

Assisted-by: ClaudeCode:claude-fable-5-1
zarr-python 2.18.7 writes `chunks: [0]` for `zarr.zeros((0,), chunks=False)`
and for `chunks=(0,)`, so stores with that document exist. Rejecting them
at open would turn a previously-readable array into an error; leaving the
0 in place read uninitialised memory after a resize. Normalize the edge to
1 with a ZarrUserWarning instead — the same grid every other "one chunk
spans the axis" spelling produces — and keep rejecting a zero edge on an
axis that has data.

Assisted-by: ClaudeCode:claude-fable-5-1
Assisted-by: ClaudeCode:claude-fable-5-1
Measured against zarr 2.18.7: `zeros((0,), chunks=False)`, `chunks=-1` and
`chunks=(0,)` all write `chunks: [0]`, after which nchunks, read, write,
append, resize and reopen-then-read every raise ZeroDivisionError. There
was never a working behaviour to preserve; normalizing the edge to 1 makes
such arrays usable for the first time. Say so in the comment and fragment
instead of claiming the stores were previously readable.

Assisted-by: ClaudeCode:claude-fable-5-1
zarr 3.2.0 and 3.2.1 classified mixed chunk specs like (2, (5, 10, 5)) as
regular and stored them as a "regular" grid whose chunk_shape contains an
edge list, while laying the chunks out as a rectilinear grid. Newer versions
accepted that metadata and failed later with an unrelated TypeError.

RegularChunkGridMetadata now rejects edge lists. When reading stored
metadata, a "regular" grid with edge lists is read as the rectilinear grid
it describes, with a warning explaining how to re-save it; if rectilinear
chunks are disabled, the error says what happened and how to enable them.

Closes zarr-developers#4374

Assisted-by: ClaudeCode:claude-opus-5
Assisted-by: ClaudeCode:claude-opus-5
@read-the-docs-community

read-the-docs-community Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

@read-the-docs-community

read-the-docs-community Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Documentation build overview

📚 zarr-indexing | 🛠️ Build #34831157 | 📁 Comparing d7ff723 against latest (1187a43)

  🔍 Preview build  

No files changed.

@codecov

codecov Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.06323% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 94.61%. Comparing base (cd1e5b3) to head (d7ff723).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
src/zarr/core/metadata/upgrades.py 98.29% 3 Missing ⚠️
src/zarr/core/chunk_grids.py 93.33% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #4375      +/-   ##
==========================================
+ Coverage   94.38%   94.61%   +0.22%     
==========================================
  Files          93       94       +1     
  Lines       13205    13518     +313     
==========================================
+ Hits        12464    12790     +326     
+ Misses        741      728      -13     
Files with missing lines Coverage Δ
src/zarr/core/_json.py 100.00% <100.00%> (ø)
src/zarr/core/array.py 98.22% <100.00%> (+0.13%) ⬆️
src/zarr/core/common.py 92.34% <100.00%> (+1.76%) ⬆️
src/zarr/core/group.py 95.30% <100.00%> (+0.09%) ⬆️
src/zarr/core/metadata/io.py 100.00% <100.00%> (ø)
src/zarr/core/metadata/v2.py 90.65% <100.00%> (+1.27%) ⬆️
src/zarr/core/metadata/v3.py 96.91% <100.00%> (+1.46%) ⬆️
src/zarr/core/chunk_grids.py 96.74% <93.33%> (-0.04%) ⬇️
src/zarr/core/metadata/upgrades.py 98.29% <98.29%> (ø)

... and 4 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

…n the 3.2.x shim

Review follow-ups for zarr-developers#4375:

- RegularChunkGridMetadata now accepts numpy integer scalars and stores
  Python ints, mirroring parse_shapelike. The strict integer check had
  reported np.int64(2) as if it were a list of chunk edges.
- The mixed "regular" grid reader converts tuple entries to lists before
  delegating to RectilinearChunkGridMetadata.from_dict, so metadata dicts
  built in Python are handled the same as parsed JSON.
- The error and warning share one message prefix.
- Tests cover numpy ints, run-length encoded edges, and tuple input, and
  the test section header says the reader is broader than the 3.2.x bug.

Assisted-by: ClaudeCode:claude-fable-5-1
…k shapes

The chunk_shapes parsers checked exact types: `from_dict` accepted a
dimension only as `int` or `list`, `expand_rle` accepted an RLE pair only
as a `list`, and the reader for 3.2.x mixed grids special-cased `tuple` to
convert it to a list before handing it to `from_dict`. A metadata dict
built in Python holds tuples where parsed JSON holds lists, and a numpy
array is as good as either, so these checks rejected valid input.

One structural predicate, `declares_chunk_edges`, now answers "is this a
sequence of edges rather than a single integer size" for all of them:
any non-integer iterable except `str`/`bytes`. It is a `TypeGuard`, not a
`TypeIs`, because it is False for strings, which are iterable. The four
sites that make that decision use it: `parse_chunk_grid`'s detection of a
mixed grid, the mixed-grid reader, `RectilinearChunkGridMetadata.from_dict`,
and `expand_rle`. The shim no longer needs its tuple special case.

Integers are `int | np.integer` throughout, matching `_parse_chunk_shape`,
and every parser stores Python ints: `_validate_chunk_shapes` now
coerces, so constructing a rectilinear grid from numpy values works the
way it already did for a regular grid.

Assisted-by: ClaudeCode:claude-opus-5
Parsing a regular chunk shape repeated its work at two levels:

- `_parse_chunk_shape` type-checked and coerced every dimension, then
  handed the result to `_validate_chunk_shapes`, which re-ran the same
  isinstance test and coercion and added only the `>= 1` check. Its
  edge-list branch was unreachable from this caller, which is why a
  `cast` was needed on the way out.
- `RegularChunkGridMetadata.from_dict` parsed the chunk shape and passed
  it to the constructor, whose `__post_init__` parsed it again.

Together that was four passes over the dimensions for `from_dict`. It
is now one: `_parse_chunk_shape` checks the range itself and no longer
calls the rectilinear validator, and `from_dict` hands the dimensions to
the constructor unparsed. Beyond the redundancy, the shared validator
was the coupling that let a rectilinear chunk shape be stored as a
regular grid (zarr-developers#4374), so the two grid kinds now validate separately.

`RectilinearChunkGridMetadata.from_dict` also had its own `>= 1` check
for bare-int dimensions, duplicating `_validate_chunk_shapes`, which
`__post_init__` runs over the result anyway. It now only puts the JSON
into shape (integer vs sequence, RLE expansion), and a bad bare int is
reported by the validator, which names the dimension. `expand_rle` keeps
its own checks because it is called directly.

Assisted-by: ClaudeCode:claude-opus-5
… format 3

The compatibility policy for a stored chunk size of 0 on a zero-length
axis covered Zarr format 2 only, so arrays written by zarr-python 3.0 and
3.1 with `chunk_shape: [0]` — or `[false]`, which 3.0 wrote for
`chunks=False` — still could not be opened at all.

`ArrayV3Metadata` now applies the same policy to a regular chunk grid: a
stored chunk size of 0 on a zero-length axis is read as 1 with a
`ZarrUserWarning`, and a zero chunk size on a positive-length axis is
left for the chunk grid parser to reject. It runs in `__init__` rather
than in the grid parser because the policy needs the array shape, which
chunk grid metadata does not carry.

Both warnings now say how to store a corrected chunk size — open the
array writable and call `array.update_attributes({})`, which rewrites the
whole document from the parsed metadata — from one shared constant. The
Zarr format 2 warning also named only zarr-python 2.x; measured against
real installs, every 3.x release before 3.4 wrote a zero chunk size for
an empty array too (3.3.0 for `chunks=-1` and `chunks=False`).

Tested against stores written by zarr 3.0.10, 3.1.6 and 3.3.0: they open,
append without losing data, and re-save to a chunk size that reopens
without a warning.

Assisted-by: ClaudeCode:claude-opus-5
d-v-b and others added 8 commits September 19, 2026 18:05
… in one routine

The legacy zero-chunk policy was written out twice: inline in
`ArrayV2Metadata.__init__`, and again in a Zarr format 3 helper. The V2
copy also zipped with `strict=False` and re-appended any trailing chunk
entries only so that a separate length check, `parse_metadata`, could
report a dimensionality mismatch after construction.

`parse_stored_chunk_shape` in `zarr.core.metadata.common` is now the one
place a stored chunk shape is checked against its array's shape, for both
formats: one entry per axis, every integer chunk size at least 1, and a
size of 0 (or JSON `false`) on a zero-length axis read as 1 with a
warning that names the writer and how to re-save. Non-integer entries,
such as edge lists, pass through for the caller's own parser.

`ArrayV2Metadata.__init__` calls it directly and `parse_metadata` is
gone. The Zarr format 3 adapter only locates a regular grid's
`chunk_shape` in the stored document and hands it over; it still runs in
`ArrayV3Metadata.__init__` because chunk grid metadata has no array shape.
The Zarr format 3 import changes that existed only for the old helper are
reverted.

Tests for the policy now target the routine: one table of valid and
legacy inputs, and one test per rejection (dimension mismatch, zero on a
non-empty axis, negative). They replace metadata-level tests in
test_v2.py and test_v3.py that only re-tested the same rules; the
end-to-end tests still cover both formats' wiring against stored arrays.

Assisted-by: ClaudeCode:claude-opus-5
…nk grids

`parse_stored_chunk_shape` passed non-integer entries through "for the
caller's own parser", which made a regular-grid policy look like a
general chunk shape routine and let it decide what a 0-length chunk means
for grids it does not own. A rectilinear grid, or any other grid, is free
to define its own semantics for 0-length chunks.

It is now `parse_stored_regular_chunk_shape`, typed `Sequence[int]`, with
no pass-through, and its docstring says it applies to Zarr format 2
`chunks` and Zarr format 3 `regular` grids only. The Zarr format 3 caller
hands it a chunk shape only when the grid is named `regular` and every
entry is an integer (`_is_regular_chunk_shape`); anything else is not a
regular chunk shape and goes to the chunk grid parser untouched.

Assisted-by: ClaudeCode:claude-opus-5
… into fix/mixed-regular-chunk-grid-4374

# Conflicts:
#	src/zarr/core/metadata/v3.py
…place

zarr-developers#4334 read a zero chunk size on an empty axis in `ArrayV3Metadata.__init__`
and zarr-developers#4375 read a 3.2.x mixed grid inside `parse_chunk_grid`, each with its
own predicate, warning text and re-save advice. Both are compatibility
readings of a stored `regular` grid, so `_read_stored_regular_chunk_grid`
now dispatches to both, `parse_chunk_grid` accepts only what the spec
allows, and both warnings use `RESAVE_METADATA_HINT`.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…panning chunk

A stored chunk size of 0 was tolerated only on a zero-length axis and
rejected otherwise, because the metadata supposedly could not say how the
stored chunks were laid out. But a chunk size of 0 gives a grid of zero
chunks, so no release could store a chunk under it, and the writers of
that metadata let the axis grow: 3.4.0 appends to a Zarr format 2 array
created empty by 3.3.0 (shape grows, no chunk written), and 3.1.6 records
a Zarr format 3 resize and the resize half of a failed append. Measured
with real installs. Those arrays open in 3.4.0, attributes included; the
rejection would have made them unopenable.

`parse_stored_regular_chunk_shape` now reads a stored 0 (or JSON `false`)
on any axis as one chunk spanning it, `max(extent, 1)`, which is what the
`-1`/`False` spec that wrote it meant. On a grown axis the warning also
says that data written to it was not saved. Negative sizes are still
rejected.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The only state machine that touched arrays compared zarr on one store with
zarr on a MemoryStore, so a chunk grid bug showed up identically on both
sides; it had no append rule, covered Zarr format 3 only, and kept every
empty axis at 0 when resizing, which is where the zero-length bugs live.

`ArrayLifecycle` checks one array against a NumPy model across both
formats, every chunk spelling (-1, False, "auto", ints, sharded,
rectilinear) and the stored chunk size of 0 that releases before 3.4
wrote, including on an axis those releases grew. Rules append, resize
(growing and shrinking to and from 0), write and re-save the metadata;
the invariant reopens the array and compares shape, values and whether
the legacy warning is due. Deliberately breaking the grown-axis policy,
the legacy warning, or append on an empty axis each fails it.

`resize` keeps partly retained chunks whole, so cells cut off by a shrink
can come back with their old values when the axis grows (as in 2.x); the
model marks such cells unknown until written instead of encoding that.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… into fix/mixed-regular-chunk-grid-4374

# Conflicts:
#	src/zarr/core/metadata/v3.py
d-v-b and others added 19 commits September 26, 2026 14:52
…ber upgrades

After the merge, storing group metadata refreshed each upgraded
consolidated member and stored its upgrade before encoding the group, and
`Group.delitem` (which encodes before deleting) skipped the refresh.

`encode_node` refreshes the consolidated members (reading only) and
encodes the group, so the rectilinear gate runs for every member before
anything is stored; `store_node` then stores the group's documents and the
members' upgrades. `save_metadata` and `Group.delitem` both use the pair,
so a group write refused without the flag stores no member upgrade either.

Tests: group attribute writes join the store-untouched table; a refused
group write stores no other member's upgrade; with the flag, deleting a
member or setting a group attribute stores the rectilinear grid in the
member and consolidated documents. Error-message expectations follow the
patch release's integer rule wording, and the mixed-grid rows expect the
JSON `true` and float edge readings to be silent.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…shape is invalid

A stored outer chunk size of 0 of a sharded array is read in multiples of the
inner chunk size. When that inner size is itself 0 or `false`, the unit is
unknown: the upgrade now leaves the 0 for the constructor, which rejects it
with the same `ValueError` zarr 3.4.0 raised, instead of a ZeroDivisionError.
`_read_codec` now matches the sharding codec like `_inner_chunk_shape` does.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pt it

A group handle kept the consolidated copies of upgraded members flagged, so
every later group write re-read every such member, one after another, and
read each twice on the first write (once to parse, once to diff).

- `save_metadata` returns the metadata it stored; `AsyncGroup._save_metadata`
  (attrs, `update_attributes`, `delitem`, `consolidate_metadata`) and
  `Group.update_attributes_async` adopt it, so later writes read no member.
- The refresh visits only nested groups and flagged members, and reads them
  with one `asyncio.gather`. The documents it reads are the ones the upsert
  diffs against (`read_documents` + `upsert_metadata(..., stored)`), and the
  array write path does the same.
- `parse_stored_array` (documents -> silently marked metadata) replaces
  `read_stored_array`'s `(metadata, bool)` tuple, and is the one "read, mark,
  don't warn" path, also for code-built metadata.
- A flagged member whose own document is gone or cannot be read (invalid,
  replaced by a group) keeps its consolidated copy instead of failing every
  group write.
- `parse_array_metadata` of a metadata object reads it as a stored document
  only when `ArrayV2Metadata.chunks` holds a 0, the one chunk size that
  constructors still accept and the upgrades change. This removes the
  per-construction `to_dict` + upgrade (AsyncArray() back to 3.4.0 speed)
  and building arrays from codec configurations holding NumPy scalars
  works again, as in 3.4.0.
- `create_hierarchy` documents that it stores a group's consolidated
  metadata as given.
- Tests: second group write reads no member; nested member refresh; member
  without a readable document; NumPy-scalar codec configuration; stale
  handle whose inner chunk shape alone changed (kills `_chunk_layout`
  returning only the outer grid).

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolutions:
- io: 4334's `read_documents`/`parse_stored_array`/concurrent refresh and
  `upsert_metadata(store_path, metadata, stored)`, with 4375's
  encode-before-store kept: refreshing reads and writes nothing, and
  returns the member upgrades (`MemberUpgrade`, carrying the documents it
  read) for `store_node` to store after the group is encoded; the stored
  members' flags are then cleared. `EncodedNode` also carries the
  metadata it encodes, which `save_metadata` returns for adoption.
  `upsert_metadata` encodes through `encode_documents` (errors name the
  node).
- `parse_stored_array` takes the array's path, so the rectilinear flag
  error on an array's first write still names it.
- A consolidated member whose own document the flag refuses is kept
  as the consolidated copy (4334's rule for unreadable members); encoding
  the group then refuses it, naming it in the consolidated metadata: the
  mixed-grid consolidate test expects that note.
- `delitem` keeps 4375's shape and adopts the metadata it stored.
- upgrades: 4375's mixed-grid branch on 4334's `units` (None when the
  inner shape is invalid).

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`delitem` on a consolidated group stored the group metadata it had
encoded before deleting, so concurrent deletions through one handle
each stored a member list that still held the others' members. The
up-front encode now only validates (metadata that cannot be stored still
fails with the store untouched); after the deletion the handle's current
metadata is stored and adopted, as zarr 3.4.0 stored it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The rectilinear flag's read note, the stale-handle write error and the
encode note now share the shape the stored-document warnings use: the
node, its path, then what was (not) done. The consolidated-member note
keeps its location qualifier ("in the consolidated metadata"), followed
by the group's note.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A group write replaced the handle's metadata with a copy holding the
consolidated members it read again, so the handle no longer shared its
consolidated dicts with subgroup handles taken earlier: an array deleted
through such a subgroup was stored again by the parent's next write, and
reappeared on reopen.

`_refresh_consolidated` now returns, along with the copy it encodes, the
members it replaced (with the consolidated dict that holds each), and
`save_metadata` writes them back into those dicts once the store succeeds,
unless the member was deleted or replaced meanwhile. `save_metadata` returns
nothing again, and the group callers keep their metadata objects. A failed
store leaves the members flagged, so a retry reads them again.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…solidated member

A group write reads each upgraded consolidated member again from its own
document, and kept the consolidated copy when that document could not be
read. That also swallowed the rectilinear chunks flag error, so a member
whose document now declares a rectilinear grid had its stale copy stored
as valid.

The flag error is now `RectilinearChunksDisabledError`, a `ValueError`,
and the refresh lets it propagate: the group write raises and stores
nothing. Unreadable documents keep their consolidated copy as before.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolution: the refreshed consolidated members are adopted in place, as on
loop/4334, and keep this branch's encode-before-store split. `MemberUpgrade`
and loop/4334's `_Refreshed` become one `RefreshedMember` (holding dict,
name, stale and current member, and the member's documents as read):
`store_node` stores the member upgrades with the group documents, then
adopts every refreshed member. `EncodedNode` drops `metadata` (nothing
adopts a copy any more) and `save_metadata` returns nothing. `_read_array`
re-raises `RectilinearChunksDisabledError`, now raised by
`_check_rectilinear_chunks_enabled` (declared next to it).

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…he flag

Storing a group's consolidated metadata reads each upgraded member again
from its own document; one read as a rectilinear chunk grid now raises the
flag error there (naming the array) instead of being kept stale. That
error now also says the group stored nothing, as the encoding error does.
`zarr.consolidate_metadata` without the flag reports the member this way.
The refresh test gains a mixed regular grid row.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`delitem` encodes the group metadata without the member before deleting
it, which reads each upgraded consolidated member again, then stored the
metadata through `save_metadata`, which read them all a second time.

It now adopts the members the first encoding read, after the deletion and
in place, and stores the metadata as it then is with those members'
upgrades (`store_node`), so the first deletion reads each member once.
The handle's consolidated metadata is no longer replaced by writes, so the
deletion pops from it directly.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…stored

Group writes (attributes, `update_attributes`, deleting a member,
`consolidate_metadata`, `create_hierarchy`) read each upgraded member of
the consolidated metadata again from its own document and stored its
upgrade. Those member writes raced concurrent deletions: two concurrent
`delitem`s on a consolidated group recreated a deleted array's metadata
document in most runs, which zarr 3.4.0 never did.

A group write now stores only the group's own documents and reads none.
Array metadata read from a document that had to be upgraded keeps that
document (`_stored_document`, set by `mark_upgraded`, replacing the
`_stored_document_upgraded` flag), and consolidated metadata stores such
a member as it was stored. Every reader of the consolidated metadata then
reads the member as upgraded again, and the array's own first chunk write
stores the upgrade of its current document (or refuses a changed chunk
grid) as before: the only write that upgrades a member document.

Removed: the refresh and adoption of consolidated members in
`save_metadata` (`_refresh_consolidated`, `_refresh_array`,
`_Refreshed`), and `Group.update_attributes_async`'s detour through
`save_metadata`. Tests pinning the refresh are replaced by tests that a
group write reads nothing and writes only the group's documents, stores
the member's consolidated copy byte for byte as stored, that concurrent
deletions leave no member, and that the first write through consolidated
metadata stores the member's upgrade or refuses a changed grid.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Group writes store only the group's own documents (D19). Resolution:

- io.py: `RefreshedMember`, `_read_array`, `_refresh_consolidated`,
  `EncodedNode`, `encode_node` and `store_node` are gone; `save_metadata`
  stores what `encode_documents` encodes.
- `AsyncGroup.delitem` encodes the group without the member (the store
  and the group stay untouched if that fails), deletes the member,
  removes it in place from the shared consolidated metadata, and stores
  the group's documents encoded from its metadata as it then is. It no
  longer reads or stores any member document.
- `GroupMetadata.to_buffer_dict` gates the consolidated members it
  encodes; a member read from a document that had to be upgraded (a
  3.2.x mixed regular grid, say) is stored as it was stored, so it is
  not gated.
- Tests: the mixed-grid group tests now pin that group writes and
  `consolidate_metadata` keep the verbatim document, with or without the
  flag; deleting a member with the flag off is refused when the group
  would store a rectilinear chunk grid.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tadata

Setting attributes of an array read from a document that had to be
upgraded stored the upgraded document with the new attributes, but left
the metadata marked with the document it was read from. A consolidated
group handle shares that metadata, so its next write stored the old
document in the consolidated metadata and the new attributes were lost
from it (zarr 3.4.0 kept them).

Every write of an array's own documents (creating, resizing, setting
attributes) now goes through `AsyncArray._save_metadata`, which clears
the mark afterwards, as the first chunk write does after storing the
upgrade. Both clear it through `_stored_document_replaced`, the one
place that declares it: the first chunk write clears it also when it
stores nothing (the store already holds a valid document, or none), so
it cannot be folded into the save.

The deep copy in `mark_upgraded` stays: the metadata's attributes share
objects with the document it was read from, which the caller also
holds. A test pins it.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ment mark

Setting attributes on or resizing an array read from a document that needed
an upgrade clears the `_stored_document` mark only after the store accepts the
new metadata. A test now injects a store failure for the array's own document
and checks the mark and the stored bytes survive, for Zarr formats 2 and 3.
The `_save_metadata` docstring now says only what the method does.

Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assisted-by: ClaudeCode:claude-opus-5-5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@d-v-b d-v-b changed the title fix: read mixed regular/rectilinear chunk grids written by zarr 3.2.x fix(chunk-grids): read mixed regular/rectilinear chunk grids written by zarr 3.2.x Sep 29, 2026
A stored chunk size of 0 or `false` was read as one chunk spanning the
axis at open time. On an axis that grew after the size was written, that
made the chunk size depend on when the array was opened, and it could be
very large: `chunks: [0, 100, 100]` on shape (10000, 100, 100) read as one
800 MB chunk. It is now read as the smallest valid chunk edge length (1,
or the inner chunk size for a shard), which is what `chunks=-1` gives on
a zero-length axis. A resize by software that kept the stored 0 no longer
changes the layout a stale handle expects.

Assisted-by: ClaudeCode:claude-opus-5-5
# Conflicts:
#	src/zarr/core/metadata/upgrades.py
#	tests/test_metadata/test_upgrades.py
…SON input as is

A stored `true` chunk size or an integral float edge is read as the value
zarr already read it as, so chunks written under the upgraded metadata are
where every reader looks. Marking such documents for re-saving turned a
plain chunk write into a metadata write: in a ZipStore that adds a second
`zarr.json` entry (a `UserWarning`, an error under `-W error`), and it
opened a race with concurrent metadata writes that 3.4.0 did not have. Each
upgrade now reports whether it moves chunks, and only those mark the
metadata.

The upgrades also detected changes by comparing JSON encodings, so
`ArrayV3Metadata.from_dict` given metadata built in code (a codec instance,
a NumPy integer in a sharding codec's `chunk_shape`) raised `TypeError`.
Changes are now tracked where they are made.

Assisted-by: ClaudeCode:claude-opus-5-5
# Conflicts:
#	src/zarr/core/metadata/io.py
#	src/zarr/core/metadata/upgrades.py
With the rectilinear flag checked when metadata is stored rather than when
it is built, `zarr.create(..., overwrite=True)` with rectilinear chunks
and the flag off deleted the existing array and then raised; zarr 3.4.0
raised before deleting anything. `AsyncArray._create_v2`, `_create_v3`
and `init_array` now build and encode the new metadata before
`_prepare_overwrite`, so metadata that cannot be stored fails with the
store untouched. This also covers `create_array(..., overwrite=True)`,
which deleted the existing array before validating its arguments.

Assisted-by: ClaudeCode:claude-opus-5-5
@d-v-b d-v-b added this to the 3.4.1 milestone Sep 29, 2026
…an overwrite test

Assisted-by: ClaudeCode:claude-opus-5-5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mixed regular/rectilinear chunk grids written by 3.2.x are unreadable by 3.3+

1 participant