Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Typesense

typesense is an output that writes documents to a Typesense collection. Its import API takes one document per line and answers with one result line per document, so a rejected document fails alone. Tested against Typesense 29.

mqb copy --drain --batch-size 10000 \
  'file://books.jsonl?format=raw' \
  'typesense://localhost:8108/books?api_key=<api key>'

typesense+https://host/books connects over HTTPS.

On this page:

Before the first run

Typesense does not create a collection on import. Create it once:

curl -X POST http://localhost:8108/collections \
  -H 'X-TYPESENSE-API-KEY: <api key>' \
  -H 'Content-Type: application/json' \
  -d '{"name": "books", "fields": [
        {"name": "title", "type": "string"},
        {"name": "year", "type": "int32"}]}'

Every document needs an id, and it must be a string. A numeric id, which is what an integer primary key gives, is rejected with “Document’s id field should be a string”. The transform below turns it into one and leaves every other field as it is.

Keep a collection in sync with a Postgres table

books_to_search:
  batch_size: 10000
  input:
    postgres_cdc:
      url: "postgres://user:pass@localhost/app"
      publication: "books_pub"
      slot_name: "mqb_typesense"
      consume: capture_all                 # backfill first, then stream changes
      cursor_id: "books_backfill"
      checkpoint_store: "file:///var/lib/mq-bridge/books-phase.json"
  output:
    middlewares:
      - transform:
          schema:
            type: object
            properties:
              id: { type: string }
      - retry: { max_attempts: 5 }
    custom:
      name: typesense
      config:
        url: "typesense://localhost:8108"
        api_key: "<api key>"
        collection: books
        operation: "${metadata:postgres.operation}"

Inserts and updates become upserts. A delete removes the document whose id is the deleted row’s primary key, so the table needs a primary key named id. See Postgres CDC for the publication and the slot.

One-off load

echo '{"type":"object","properties":{"id":{"type":"string"}}}' > id-as-string.json

mqb copy --drain --batch-size 10000 \
  'file://books.jsonl?format=raw' \
  'typesense://localhost:8108/books?api_key=<api key>|transform?schema_file=id-as-string.json'

Loading 20,000 small documents this way took under a second on a laptop.

Options

FieldDefaultMeaning
urlrequiredtypesense://host:8108, typesense+https://host or an http(s):// URL
collectionrequiredCollection to write to; the path of the URI
api_keynoneSent as X-TYPESENSE-API-KEY
operationnoneTemplate for a message’s operation; delete or d removes the document
compressionnonegzip, zstd or lz4 request bodies
request_timeout_msnoneRequest timeout

typesense is the generic http_bulk output with these requests filled in. For anything it does not offer, such as another import action or extra headers, write the http_bulk form:

output:
  http_bulk:
    url: http://localhost:8108
    headers:
      X-TYPESENSE-API-KEY: <api key>
    operation: "${metadata:postgres.operation}"
    upsert:
      path: /collections/books/documents/import?action=upsert
      content_type: text/plain
      result:
        lines:
          success: /success
          error: /error
    delete:
      method: DELETE
      path: /collections/books/documents?filter_by=id:[{ids}]
      max_ids: 100

What to know

  • A document without id is still written. Typesense invents an id, and a later delete cannot find the document.
  • A rejected document fails alone, with Typesense’s reason, for example “Field year must be an int32”. Add a dlq to keep it.
  • Deletes go into the URL, 100 ids per request, and an id containing a comma cannot be deleted this way.
  • No collection management. Schema changes, aliases and synonyms stay with the Typesense API.

The behaviour on errors is described on the http_bulk page.