Skip to content

dataset loader improvements - #737

Open
jshook wants to merge 9 commits into
mainfrom
jshook_ashwin/ds-loader-improvements
Open

jshook wants to merge 9 commits into
mainfrom
jshook_ashwin/ds-loader-improvements

Conversation

@jshook

@jshook jshook commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Thanks to @ashkrisk for the collab on this. We had a discussion about how to tackle the needed changes in the loader for various testing requirements and dataset sizes, and this PR is what came out of it, as an idea to tackle all the concerns we both raised.

What it does

  • Programmatically remove any in-built notion of a file-based vector data slab which is simply loaded into memory.
    • This is done by removing the extant DataSet.getBaseVectors() list accessor; base vectors are only reachable through getBaseRavv(), and contiguous subsets use the new RandomAccessVectorValues.range(from, to) view instead of subList.
  • Provide a path to having in-memory datasets, with a proper modular behavior, via DataSetWrapper. The loader memory-maps the fvecs file; wrappers decide where the base vectors live at runtime:
    • memory – heap-resident copy (the default, so existing callers behave as before)
    • mmap – served from the mapped file
    • lru – bounded grain-resolution LRU cache in front of the mapped file, for larger-than-memory datasets (options grain, capacityMb)
  • Dataset selection carries a loader profile and a wrapper list, as name:profile(wrapper[k=v],...) or as a structured datasets.yml entry (name, profile, wrappers).
    • Catalog keys may be name:profile; loaders that don't understand profiles accept default and reject anything else.

Also fixed along the way: Grid, CompactorBenchmark and ParallelWriteExample read the base reader directly from parallel writer threads. With a value-shared (mapped) reader this corrupted inline/NVQ features (siftsmall recall 0.85 vs 0.99). They now go through threadLocalSupplier().

Verification:

  • New unit tests for the range view, mapped reader, spec parsing, wrappers, LRU cache, partitioner and collection parsing; existing loader tests updated.
  • 3.9 GB dataset: mapped load + heap cache 0.2–0.8 s vs 4.2 s for the old streaming reader; full scan through mmap and lru under a 1 GB heap.
  • BenchYAML end to end on siftsmall through all three wrappers and a profile entry, on both the array and native providers; recall unchanged.
  • Each commit compiles standalone.
  • Docs: docs/benchmarking.md covers the new dataset syntax.

Introduce range(fromOrdinal, toOrdinal) as a default method returning a
re-based, non-copying view, implemented by RangeRandomAccessVectorValues.
Nested ranges collapse onto the backing reader, the view's size is fixed at
creation, and value-sharing and copy semantics follow the backing reader.
This lets consumers address contiguous subsets of a dataset without the
list-structured accessor.
MappedFvecsRandomAccessVectorValues maps an fvecs file read-only in slabs of
at most 1 GiB, so files beyond 2 GB work, and serves vectors by copying the
record out of the mapping with a per-record dimension check. It is
value-shared; copies share the mapping and own a scratch vector. SiftLoader
gains writeFvecs so any reader can be spilled to the same format.
…wrappers

Remove DataSet.getBaseVectors so base vectors are reached only via
getBaseRavv, and introduce DataSetWrapper to decide where they live.
DataSetWrapper layers over an origin DataSet and delegates every accessor
to it; Provider builds a wrapper from an origin and Factory builds a
Provider from symbolic options. InMemoryCachedDataSet reads the origin's
base vectors through ranged views in parallel chunks into heap memory,
adopting list-backed origins as they are. MMapCachedDataSet serves base
vectors from a memory-mapped fvecs file, adopting mapped origins and
spilling heap-resident ones to a cache file.

SimpleDataSet accepts any reader, the multi-file loader maps the base fvecs
file instead of reading it into a list, and DataSets applies wrapper
providers to what loaders return, caching base vectors in heap memory by
default so existing callers keep their previous behaviour. Legacy scrubbing
copies the mapped vectors into heap first, since it must rewrite them.
DataSetInfo now delegates loadBehavior.

Consumers migrate to the reader: DataSetPartitioner hands out ranged views,
the compaction bench and the JMH compactor benchmark build and search from
readers directly, and the remaining call sites only needed sizes.
FvecsLoadEconomyTest compares the previous streaming loader with the mapped
reader plus in-memory cache.
LruGrainRandomAccessVectorValues keeps a bounded number of grains of
consecutive vectors in heap memory in front of any reader, loading a grain
through a ranged view on a miss and evicting the least recently used grain
when the bound is exceeded. Reads hit a concurrent map without locking,
distinct grains load in parallel, concurrent misses on one grain wait for
the first loader, and a failed load is withdrawn so the next reader retries.
LruCachedDataSet wraps it with grain and capacity settings that default to
system properties or a quarter of the heap. LargerThanHeapDataSetTest scans
a file through the mmap and lru wrappers under a heap smaller than the file.
DataSetSpec names a dataset with an optional loader profile and an ordered
list of wrappers, each with options. The sugared form is name,
name:profile, name(wrapper,...) or name:profile(wrapper,...), where a
wrapper may carry options as lru[grain=4096,capacityMb=512]; the structured
YAML form is a map with name, profile and wrappers, where a wrapper entry is
a name, a map keyed by the wrapper name holding its options, or a map with a
name key and options alongside. toString renders the canonical sugared form
losslessly, so DatasetCollection can keep passing strings through the
existing name-based filtering and configuration pipeline.

DataSets resolves symbolic wrappers through a registry of factories that
receive each wrapper's options; memory and mmap take none, lru accepts grain
and capacityMb. A missing profile is reported as default. Loaders that do
not understand profiles accept default and reject any other profile for a
dataset they have, while still returning empty for names they do not
recognise. MultiConfig looks up per-dataset config files by the bare name.
Catalog entries may be keyed as name:profile, as in sift1m:label_00, so
one dataset can ship several variants of its files. Loading by DataSetSpec
now resolves against those keys: the default profile matches the bare name
entry and then name:default, while any other profile matches only
name:profile and is otherwise simply not found. Loading by string still
treats the argument as a literal key. Metadata is looked up by the matched
key and then by the bare name, through a new multi-key lookup on
DataSetMetadataReader that names the result after the requested dataset,
so profile variants can share one metadata entry while keeping their full
name in results. The memory and mmap wrappers now log when they adopt an
origin that is already in their form.
…suppliers

The feature-state suppliers that Grid, the JMH compactor benchmark and the
parallel-write example hand to on-disk writers are invoked from several
writer threads, yet each read the dataset's base reader directly. That is
fine for a list-backed reader but not for a value-shared one such as the
memory-mapped fvecs reader, whose getVector returns a per-instance scratch
vector: concurrent writer threads overwrote each other's vectors, so inline
and NVQ features were written from corrupted data and recall on a dataset
selected with the mmap wrapper dropped from 0.99 to 0.85 on siftsmall. The
suppliers now read through RandomAccessVectorValues.threadLocalSupplier,
which is a no-op for un-shared readers and restores parity for shared ones.
FvecsLoadEconomyTest and LargerThanHeapDataSetTest print which
VectorizationProvider is in effect, so a run's output states whether the
array or the native provider was exercised.
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Before you submit for review:

  • Does your PR follow guidelines from CONTRIBUTIONS.md?
  • Did you summarize what this PR does clearly and concisely?
  • Did you include performance data for changes which may be performance impacting?
  • Did you include useful docs for any user-facing changes or features?
  • Did you include useful javadocs for developer oriented changes, explaining new concepts or key changes?
  • Did you rebase your branch onto the latest main for regression testing and PR submission?
  • Did you trigger regression testing via Run Bench Main and review results?
  • Did you adhere to the code formatting guidelines (TBD)
  • Did you group your changes for easy review, providing meaningful descriptions for each commit?
  • Did you ensure that all files contain the correct copyright header?
  • Did you add documentation for this feature to the release notes directory?

If you did not complete any of these, then please explain below.

Windows refuses to modify a file that has a live memory mapping, which the
examples tests hit on the JDK 20 Windows job: a second mmap selection of the
same dataset rewrote a spill file the first one still mapped, and the fvecs
round-trip test rewrote a file it had just mapped. MMapCachedDataSet now
reuses an existing spill file whose size matches, and otherwise writes to a
temporary file beside it and moves it into place, so a mapped file is never
opened for writing. Reuse also makes the spill a one-time cost per dataset
rather than per run. The round-trip test proves overwrite on an unmapped
file, and the wrapper tests cover reuse of a mapped spill and replacement
of a stale one.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant