Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Directory spool

Stores each message as a payload file in a directory, with an optional JSON metadata sidecar. Use it as a durable FIFO hand-off between processes when you want a queue that can be inspected with ordinary filesystem tools and do not want to operate a broker.

Unlike the file connector and local object storage, which can frame many records in a file, the directory spool writes one opaque payload per file. A producer may exit while a consumer continues draining the backlog.

URL format

dir-spool:///absolute/path/to/spool?<option>=<value>

The aliases spool:// and dirspool:// are also accepted. The directory path comes from the URI path, not a path query parameter.

Config (YAML / library)

The same settings as a route endpoint in a config file, or in Route.from_config / fromConfig. Every URL query parameter is a field of the same name under dir_spool:.

output:
  dir_spool: { path: "/var/spool/orders" }

The URL path becomes the path field.

Source, target, and on-disk layout

As a target (sink), the connector writes each incoming message as one chunk:

/var/spool/orders/
├── 000000000.bin   # raw payload bytes for message 0
├── 000000000.json  # message id and string metadata for message 0
├── 000000001.bin   # raw payload bytes for message 1
├── 000000001.json  # message id and string metadata for message 1
├── PRODUCER        # producer claim file
├── CONSUMER        # draining-consumer claim file, when one is running
└── DONE            # optional completion sentinel

The default {seq:09} naming pattern produces the zero-padded names above. payload_extension and metadata_extension change only the suffixes. A sidecar has this shape:

{
  "message_id": "0195f4a0c7d77801a2b3c4d5e6f78901",
  "metadata": {
    "content_type": "application/json",
    "source": "orders"
  }
}

The payload file contains exactly the message payload, byte for byte. It may be UTF-8 text, JSON, an image, compressed data, or any other binary content. The connector does not add a delimiter, envelope, or encoding. Set metadata_extension to an empty string when sidecars are not needed.

With the default atomic=true, files are first written to sibling .tmp names. The sidecar is finalized before the payload; the payload’s final rename makes the chunk visible to readers without exposing incomplete data.

As a source, the connector scans for files with payload_extension, reads each whole file as one message, and restores its message id and metadata from the matching sidecar when present. A payload without a sidecar is valid and is delivered with empty metadata, so another program can produce chunks without knowing the sidecar format. Such a producer should likewise finalize the payload last.

With sharding enabled, the leading sequence digits become directories. For example, {seq:09}, shard_depth=2, and shard_width=3 store message 1 as:

/var/spool/orders/000/000/001.bin
/var/spool/orders/000/000/001.json

Both source and target must use the same sharding depths, widths, and file extensions.

Payload formats and batching

dir_spool does not parse or generate JSONL, CSV, or other record-oriented formats. A .jsonl or .csv payload file is still one opaque message, even if it contains many lines. Changing payload_extension changes file selection and naming, not parsing behavior.

Use a record-oriented connector when records must be framed:

  • The file connector reads or writes one named file, for example file:///data/orders.jsonl?format=json or file:///data/orders.csv?format=csv.
  • Local object storage watches a directory of immutable, multi-record files, making it a better fit for CSV/JSONL ETL drop zones.
  • Sending either source to a directory spool creates one chunk per emitted message, rather than one chunk containing the whole input file.

Route-level batching is supported in both directions. A target can receive a batch of messages and writes one payload/sidecar pair for each message. A source can return up to the route’s batch-size limit per read. Batching therefore reduces routing and filesystem-sync overhead but does not change the one-file, one-message chunk format.

Examples

Write a finite input to a spool and mark it complete:

mqb copy --drain \
  --from file:///data/orders.jsonl?format=json \
  --to 'dir-spool:///var/spool/orders?emit_done=success'

Drain that spool into PostgreSQL, then exit:

mqb copy --drain \
  --from 'dir-spool:///var/spool/orders?stop_on_done=true' \
  --to 'postgres://user:pass@localhost/app?table=orders'

This one-shot example starts after the producer command has completed, so --drain exits when the backlog is empty. In a continuously running route, stop_on_done=true ends the source only when both the queue is empty and the producer’s DONE sentinel exists. Set emit_done only on the last producer; a producer opening the spool removes a stale sentinel before writing again.

Shard a high-volume spool:

mqb copy \
  --from mqtt://broker.local:1883?topic=telemetry \
  --to 'dir-spool:///var/spool/telemetry?naming_pattern={seq:012}&shard_depth=2&shard_width=3'

Configure the consumer with the same shard_depth, shard_width, payload extension, and metadata extension. Sharding uses leading sequence digits as subdirectories and avoids placing an unbounded number of files in one directory.

Delivery and concurrency

Chunks are delivered in lexical filename order. Keep a zero-padded sequence at the start of naming_pattern; the default {seq:09} is safe. The producer writes temporary files and renames them into place, so the consumer does not observe a partial chunk. With the default fsync=chunk, acknowledged writes are also flushed to durable storage.

A draining consumer deletes a chunk only after its message is acknowledged. A nack leaves the chunk on disk for redelivery. Set drain_on_read=false for a non-destructive pass that leaves files in place.

By default, one producer and one draining consumer may use a spool at the same time. Separate PRODUCER and CONSUMER lock files enforce those roles without preventing the normal producer-consumer pair. Avoid disabling these claims: multiple producers can collide, and multiple draining consumers can deliver duplicates because claiming is in memory rather than an on-disk rename.

Key options

OptionPurpose
naming_patternSink only: chunk name template. It must begin with a sequence; default {seq:09}.
payload_extension / metadata_extensionFile suffixes. Set metadata_extension to an empty string to omit sidecars.
atomicSink only: write through a temporary file and rename; default true.
fsyncchunk (default) for durable writes, or off for higher throughput with weaker crash guarantees.
emit_doneSink only: write the completion sentinel on success, on any end, or never (default).
stop_on_doneSource only: exit after the sentinel is present and the backlog is empty.
drain_on_readSource only: delete acknowledged chunks; default true.
shard_depth / shard_widthSpread chunks across sequence-derived subdirectories. Both ends must agree.
claimexclusive (default), warn, or off for same-role locking.

Full field list: reference/dir-spool.md.

For detailed durability, sharding, lock, and completion-sentinel behavior, see the configuration guide.