Repository navigation
Streaming uploads and downloads across storage clients and storages #2240
Description
Activity
- addedt-toolingIssues with this label are in the ownership of the tooling team.Issues with this label are in the ownership of the tooling team.
on Sep 16, 2026 What happens today
Uploads
KeyValueStore.set_value(src/crawlee/storages/_key_value_store.py:173) forwardsvalue: AnytoKeyValueStoreClient.set_value(src/crawlee/storage_clients/_base/_key_value_store_client.py:53). Neither says whatvaluemay be, and the backends disagree:- File system (
_file_system/_key_value_store_client.py:302), SQL (_sql/_key_value_store_client.py:149) and Redis (_redis/_key_value_store_client.py:125) serialize through a chain that ends instr(value).encode('utf-8'). A file handle or a generator is stored as itsrepr, soawait kvs.set_value('out', open('big.bin', 'rb'))writes the literal text<_io.BufferedReader name='big.bin'>and reports success. - Memory (
_memory/_key_value_store_client.py:112) keeps the object as it is and recordssys.getsizeof(value)as the size. A generator is then consumed by whichever reader gets to the record first. - Apify (apify-sdk-python,
src/apify/storage_clients/_apify/_key_value_store_client.py:123) hands the value straight toset_record, so once feat: Stream request bodies from files, iterables, and responses apify-client-python#1060 lands this one backend streams correctly by accident, with nothing in Crawlee's types to promise it.
So the same call is a working streaming upload on the platform and silent data loss on the local default.
Downloads
Every backend materializes the whole record. The file system client does
record_path.read_bytes(), SQL selects the full BLOB, Redis doeshget, and the Apify client callsget_recordeven thoughstream_recordis available.KeyValueStore.get_valuereturns a decoded value and there is no way to ask for a stream, so reading a 2 GB screenshot or a large NDJSON export costs 2 GB of RSS.Datasets
Dataset.export_to(src/crawlee/storages/_dataset.py:362) collects every item into a list (src/crawlee/_utils/file.py:193), serializes it into aStringIO, then passes the resultingstrtoset_value. A large dataset is held in memory at least twice before the first byte is written. This is the most obvious consumer of a streamingset_value, since the export is already produced incrementally.On the read side,
Dataset.iterate_itemsalready pages, so the gap there is smaller. Worth checking whetherget_datawith the defaultlimit=999_999_999_999deserves the same treatment.Questions to settle
- Where does the contract live? Does
KeyValueStoreClient.set_valuewiden itsvaluetype to include file-likes and (async) iterators, or does streaming get its own method (stream_value/set_value_stream) that backends opt into? The first keeps one call site, the second keepsAnyfrom meaning "anything, results may vary". - What do non-streaming backends do? SQL and Redis store one value per row/field, so they cannot stream a write in any useful sense. Buffering the source into
bytesis correct and at least well defined, unlike today'srepr. The file system backend can stream a chunked write for real. - What shape does the download take? An
AsyncIterator[bytes], an async context manager wrapping the response, or a file-like object. It has to work for a local file, a Redis field, and an HTTP response, and it has to make the Apify backend'sstream_recordusable through the storage. - What happens to the record metadata?
infer_mime_typecannot inspect a stream, so the content type has to default toapplication/octet-streamor be required from the caller, andsizeis unknown until the source is drained. - Retries and seekability. apify-client-python retries only seekable sources, rewinding before each attempt, and gives everything else a single attempt. Crawlee's backends need a rule of their own for a half-written file or a partially consumed iterator.
- Should
Dataset.export_tostream into the KVS onceset_valuesupports it? That removes theStringIOand the fully materialized item list. - Does anything on the JS side need to match?
KeyValueStore.setValuethere already accepts a stream, while reading one means dropping down to the resource client (Better KVS streaming support crawlee#2929).
✍️ Drafted by Claude Code
- File system (
- addedsolutioningThe issue is not being implemented but only analyzed and planned.The issue is not being implemented but only analyzed and planned.
on Sep 16, 2026
apify-client-python is adding streaming request bodies in apify/apify-client-python#1060. After it lands,
set_recordaccepts a file-like object, an iterator ofbytes/strchunks, or a streamedHttpResponse, and sends it chunked without buffering. The download direction (stream_record) has been there for a long time.Crawlee has no equivalent at any layer: not in the
KeyValueStoreClient/DatasetClientcontracts, not in theKeyValueStore/Datasetfrontends, and not in the backends. Passing a file object through today's API also corrupts the record silently on three of the five KVS backends.#1931 asks for streaming KVS records and is still in solutioning. This issue is the wider investigation it needs: what the interfaces should look like, what each backend can actually do, and what the storages should expose.
Related
✍️ Drafted by Claude Code