Repository navigation
Proposal: Blob.from() for creating virtual Blobs with custom backing storage #209
Description
Activity
A fundamental aspect of Blobs is that they are immutable when they are minted and do not require running content JS after they are minted. This allows them to be postMessaged between agents and origins without the recipient having to worry about whether the Blob is still usable (modulo storage clearing), as well as allowing APIs like IndexedDB to know they can reliably store a Blob without indefinitely hanging a transaction.
What you are proposing sounds like a combination of:
- A means of working around the lack of a download API or expansion of srcObject to other tags so we might be able to directly attach Response objects to an HTML tag that could initiate a download.
- A means of reproducing the ability of ServiceWorkers to create synthetic responses to network requests but without involving ServiceWorkers (but instead involving URL.createObjectURL) and requiring that streams are returned instead of Response objects.
I agree that the underlying use case of downloads without involving ServiceWorkers is poorly addressed and that existing workarounds do already frequently involve URL.createObjectURL in ways that are suboptimal from a resource-utilization perspective. I also understand that you are explicitly talking about supporting third-party storage which is a situation where it is arguably much uglier and/or more concerning for a site to need to load potentially multiple third-party JS libs into the ServiceWorker to rely on its semantics. (Foreign-fetch ServiceWorkers would have provided a means for this but they were explicitly ruled out.)
I think would be much less troubling in terms of impact on the web platform to pursue further progress on a download API and/or srcObject rather than changing the existing Blob invariants.
I do think further discussion of this proposal needs an explanation of how Blobs that are postMessaged are dealt with.
I do think further discussion of this proposal needs an explanation of how Blobs that are postMessaged are dealt with.
Yes, i know - i figured that when i wrote this. Sending a
postMessagewith a custom backing store would not be easily serializable, perhaps maybe there could be an internal flag saying that they are not structured cloneable. perhaps maybe the clone could become more like a shortcut / symbolic link so that if you try to read blob with custom storage on another thread, then it would ask the owners thread to supply data back.
One solution to it could be to send a newMessageChannelor a transferableReadableStreamof some sort when cloning it.
The other thread would just ask the owner: "Please give me the data from A to B and give me back a ReadableStream"The idea where not just only meant for downloading things, i have had this idea for a long time and also being able to create own blob's in Deno to create Blob/Files backed up by the filesystem that lacks something similar to Node's
fs.openAsBlob()i mean, i tried doing something like
Deno.openAsBlob()2 days ago: denoland/deno#27664 (comment)
it feels rather hacky - and this solution won't work in browser, cuz browser read data more privately from internal [[properties]]. backend servers rely more on public properties / methods that can be patched. often via Symbols via some hack-ish method
i could also imagine creating a objectURL of an blob with custom backing store where you assign it to a
<video>seeking in it would be easy to do using 206 partial responses.
I once created a zip reader that reads the central directory at the very end of the file to get all the entries and return them as File-Like objects
if a file is uncompressed it would be as easy to just slice a zip file of where it start / ends, but if it is compressed then it would all be async.
i first created this Http-File-Like: transcend-io/conflux#31 a very long time ago and others have done similar things
https://www.npmjs.com/package/vinyl - a very large, old popular concept of a very simple metadata object that describes a file, it's essentially just the same thing as a
Blobwith an own custom backing store. I just think it would be cool to standardize them to a actual Blob object such that they can also work together with- URL.createObjectURL()
- FileReader
- Response()
- Request()
- FormData.append()
- fetch(blob:url),
- Blob()
- File()
- createImageBitmap
- etc
Speaking of IndexedDB and transaction.
I have seen folks using IndexeDB to store large AI model into IndexedDB as blobs.
if you think about how then need to go about storing it they would 1) first have to download everything into memory 2) and only then can they store that blob.if you could create a very simple blob with own custom backing store then you would never have to fill up your memory. the file can just be piped directly to the disc
it's a win win.
sure the transaction would stall for a bit, but i think this solution is actually a good way to pipe data directly into IndexedDB without filling the memory.but ofc there is now better ways to store it using
Cachestorage or opfs. But the idea of piping data into indexeddb is still a good argument rather than creating a blob in memory first and then save it to indexeddb.I have a use case where I'm streaming through a multipart form, including a Blob. The size is not known a priori. This is OK, because the parts each have a boundary, so the size need not be known in order to be able to stream the form back out. I use Node.js' implementation of FormData with a custom
Fileclass to create the form, and I stream it back out usingFormDataEncoder. Clearly such blobs are not easily transferable, however, if we consider that they are ultimately backed by a file descriptor from which the process reads, then a mechanism can be constructed for transferring the file descriptor between processes either directly, if the platform supports this (like UNIXsendmsg()), or by having one process co-operate with the other over a named pipe.I suppose one alternative would be for the FormData API to accept
ReadableStreams as values.Blob tries hard to be a snapshot rather than a mutable object: https://w3c.github.io/FileAPI/#blob-section
Each Blob must have an internal snapshot state, which must be initially set to the state of the underlying storage, if any such underlying storage exists. Further normative definition of snapshot state can be found for Files.
Changing this may have various unexpected impact to the consumers of blobs, as @asutherland explained earlier.
Maybe we could have something like a readable version of FileSystemWritableFileStream, which will solve the issue of postMessaging better (as streams are already postMessage-able).
I'd like to propose an addition to the Blob API to enable the creation of virtual Blob or File instances backed by custom-defined storage logic. This would allow SDKs and libraries to expose file-like objects without requiring the developer to manage low-level data fetching or streaming themselves.
✨ Proposed API
size: Total size of the blob (required).
stream(start, end): Required method that returns a ReadableStream or AsyncIterable for the requested byte range. it may or may not be asyncThis API is synchronous to create, but lazy in that no data is fetched until actually needed. The internal Blob machinery would take care of slicing and offsetting, so the developers only need to focus on implement the backing source logic.
🧩 Example Usage with SDK
This enables a clean integration pattern with APIs like Dropbox, Google Drive, or internal systems:
What dropbox sdk then actually dose:
In this model:
fileHandlecontains only metadata (filename, type, size, lastModified).openAsFile()constructs a File backed by virtual Blob part..arrayBuffer().✅ Benefits
🔧 Comparison to Today
Without this feature, developers must manually wrap streams, manage slicing, and emulate Blob behavior — often with duplicated effort and edge-case bugs. This addition would make such patterns first-class citizens in the platform.
🔍 Real-World Problem This Solves
A common pattern on the web is to trigger file downloads using a Blob URL and a programmatically clicked
<a>tag:However, to use this pattern today, you must already have all the file data in memory.
If you're working with a remote file (e.g. from cloud storage or an SDK), you can't delay the download until after the click — because:
Once the click handler ends,
isTrustedbecomesfalse.Any async operation (like fetching the file) that happens after the click ends is now treated as not user-initiated.
Browsers will block the download, thinking it's an automatic or malicious attempt.
With
Blob.from(), we could instead return a virtual Blob that doesn't require any data until it's needed:In this model, the download is triggered immediately within the
clickevent — but data is only fetched as it's needed, safely within the user gesture's trusted scope.✅ This allows you to: