Conversation
Introduce range(fromOrdinal, toOrdinal) as a default method returning a re-based, non-copying view, implemented by RangeRandomAccessVectorValues. Nested ranges collapse onto the backing reader, the view's size is fixed at creation, and value-sharing and copy semantics follow the backing reader. This lets consumers address contiguous subsets of a dataset without the list-structured accessor.
MappedFvecsRandomAccessVectorValues maps an fvecs file read-only in slabs of at most 1 GiB, so files beyond 2 GB work, and serves vectors by copying the record out of the mapping with a per-record dimension check. It is value-shared; copies share the mapping and own a scratch vector. SiftLoader gains writeFvecs so any reader can be spilled to the same format.
…wrappers Remove DataSet.getBaseVectors so base vectors are reached only via getBaseRavv, and introduce DataSetWrapper to decide where they live. DataSetWrapper layers over an origin DataSet and delegates every accessor to it; Provider builds a wrapper from an origin and Factory builds a Provider from symbolic options. InMemoryCachedDataSet reads the origin's base vectors through ranged views in parallel chunks into heap memory, adopting list-backed origins as they are. MMapCachedDataSet serves base vectors from a memory-mapped fvecs file, adopting mapped origins and spilling heap-resident ones to a cache file. SimpleDataSet accepts any reader, the multi-file loader maps the base fvecs file instead of reading it into a list, and DataSets applies wrapper providers to what loaders return, caching base vectors in heap memory by default so existing callers keep their previous behaviour. Legacy scrubbing copies the mapped vectors into heap first, since it must rewrite them. DataSetInfo now delegates loadBehavior. Consumers migrate to the reader: DataSetPartitioner hands out ranged views, the compaction bench and the JMH compactor benchmark build and search from readers directly, and the remaining call sites only needed sizes. FvecsLoadEconomyTest compares the previous streaming loader with the mapped reader plus in-memory cache.
LruGrainRandomAccessVectorValues keeps a bounded number of grains of consecutive vectors in heap memory in front of any reader, loading a grain through a ranged view on a miss and evicting the least recently used grain when the bound is exceeded. Reads hit a concurrent map without locking, distinct grains load in parallel, concurrent misses on one grain wait for the first loader, and a failed load is withdrawn so the next reader retries. LruCachedDataSet wraps it with grain and capacity settings that default to system properties or a quarter of the heap. LargerThanHeapDataSetTest scans a file through the mmap and lru wrappers under a heap smaller than the file.
DataSetSpec names a dataset with an optional loader profile and an ordered list of wrappers, each with options. The sugared form is name, name:profile, name(wrapper,...) or name:profile(wrapper,...), where a wrapper may carry options as lru[grain=4096,capacityMb=512]; the structured YAML form is a map with name, profile and wrappers, where a wrapper entry is a name, a map keyed by the wrapper name holding its options, or a map with a name key and options alongside. toString renders the canonical sugared form losslessly, so DatasetCollection can keep passing strings through the existing name-based filtering and configuration pipeline. DataSets resolves symbolic wrappers through a registry of factories that receive each wrapper's options; memory and mmap take none, lru accepts grain and capacityMb. A missing profile is reported as default. Loaders that do not understand profiles accept default and reject any other profile for a dataset they have, while still returning empty for names they do not recognise. MultiConfig looks up per-dataset config files by the bare name.
Catalog entries may be keyed as name:profile, as in sift1m:label_00, so one dataset can ship several variants of its files. Loading by DataSetSpec now resolves against those keys: the default profile matches the bare name entry and then name:default, while any other profile matches only name:profile and is otherwise simply not found. Loading by string still treats the argument as a literal key. Metadata is looked up by the matched key and then by the bare name, through a new multi-key lookup on DataSetMetadataReader that names the result after the requested dataset, so profile variants can share one metadata entry while keeping their full name in results. The memory and mmap wrappers now log when they adopt an origin that is already in their form.
…suppliers The feature-state suppliers that Grid, the JMH compactor benchmark and the parallel-write example hand to on-disk writers are invoked from several writer threads, yet each read the dataset's base reader directly. That is fine for a list-backed reader but not for a value-shared one such as the memory-mapped fvecs reader, whose getVector returns a per-instance scratch vector: concurrent writer threads overwrote each other's vectors, so inline and NVQ features were written from corrupted data and recall on a dataset selected with the mmap wrapper dropped from 0.99 to 0.85 on siftsmall. The suppliers now read through RandomAccessVectorValues.threadLocalSupplier, which is a no-op for un-shared readers and restores parity for shared ones.
FvecsLoadEconomyTest and LargerThanHeapDataSetTest print which VectorizationProvider is in effect, so a run's output states whether the array or the native provider was exercised.
jshook
requested review from
MarkWolters,
ashkrisk and
tlwillke
as code owners
September 24, 2026 22:04
Contributor
|
Before you submit for review:
If you did not complete any of these, then please explain below. |
Windows refuses to modify a file that has a live memory mapping, which the examples tests hit on the JDK 20 Windows job: a second mmap selection of the same dataset rewrote a spill file the first one still mapped, and the fvecs round-trip test rewrote a file it had just mapped. MMapCachedDataSet now reuses an existing spill file whose size matches, and otherwise writes to a temporary file beside it and moves it into place, so a mapped file is never opened for writing. Reuse also makes the spill a one-time cost per dataset rather than per run. The round-trip test proves overwrite on an unmapped file, and the wrapper tests cover reuse of a mapped spill and replacement of a stale one.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Thanks to @ashkrisk for the collab on this. We had a discussion about how to tackle the needed changes in the loader for various testing requirements and dataset sizes, and this PR is what came out of it, as an idea to tackle all the concerns we both raised.
What it does
Also fixed along the way: Grid, CompactorBenchmark and ParallelWriteExample read the base reader directly from parallel writer threads. With a value-shared (mapped) reader this corrupted inline/NVQ features (siftsmall recall 0.85 vs 0.99). They now go through threadLocalSupplier().
Verification: