Skip to main content
Register a MongoDB database and its collections become a live, read-only catalog in the lake. Query them directly for ad hoc work, or import a collection into Iceberg when you want a snapshot you can run large jobs against. Connections are managed in Settings → Lake.
Point oleander at a secondary or a dedicated read-only user, and make sure your cluster’s network access list allows oleander’s egress.

Setup

  1. Open Settings → Lake and choose Add MongoDB connection.
  2. Paste a mongodb:// or mongodb+srv:// connection string to fill the fields automatically, or enter them by hand:
The dialog includes the shell commands for a read-only user if you want to create one before saving.

Querying live

Collections are reachable as connection_name.database.collection:
Nested documents are flattened for SQL: a field at address.city is the column address_city. In the lake tree, field lists are sampled from the first 200 documents and _id is shown as the key.
A MongoDB collection selects DuckDB. DuckDB is the only engine that attaches external connections, so the query router routes any query naming a connection collection there before it considers input size - leave engine on auto. Asking for Bloom, Polars, or Spark on one of these queries returns an engine capability error rather than rerouting silently.
Live reads run against your database, so filter and limit the way you would against production. Once a collection is imported it lives in Iceberg and the router is free to send it to any engine.

Importing into Iceberg

For anything larger than an ad hoc read, import the collection. The import runs as a Spark job using the MongoDB Spark connector and lands the data as an Iceberg table in your lake, where every engine can reach it. The schema is inferred from a sample of documents:
  • Nested documents are kept as structs and arrays as arrays, so the imported table is not flattened the way live queries are.
  • A field seen with conflicting types is stored as a string. A field that is only ever null is also stored as a string.
  • Integers and decimals are widened so later, larger values still fit.
  • _id is stored as a string, with ObjectIds as their 24-character hex.
The destination defaults to oleander.default.<collection> - the sanitized collection name, with no connection prefix. OVERWRITE replaces it, APPEND adds to it. An import does not return rows. It submits a job and returns a run_id you can monitor like any other run, and once it completes you can sample the destination with a normal query.

Collection syncing

A one-off import gives you a snapshot. Put the same import on a schedule and the Iceberg copy tracks the source instead. Schedules run hourly or daily, and are managed the same way as Postgres syncs: under Scheduled syncs in Settings → Lake, and under Transfers in the lake’s left panel.
Scheduled syncs are incremental. After the first full import, each run merges new and changed documents on _id rather than recopying the collection.
How a run finds what changed depends on the collection, and the reason is written to the run’s plan log:
  • A timestamp field such as updatedAt, updated_at, lastModifiedAt, or modifiedAt means documents modified since the last run are merged.
  • Otherwise, if _id values are ObjectIds, documents inserted since the newest one’s embedded creation time are merged. Updates to existing documents are not picked up on this path.
  • Anything else falls back to a full refresh.
New top-level fields that appear in the source are added as columns; new fields inside an existing struct wait for the next full refresh. Deletes are never synced. A document removed from MongoDB stays in the Iceberg copy. Every sync run is a normal oleander run: it appears in run history, emits lineage with the collection as the input and the Iceberg table as the output, and carries cost like any other job.

From an agent

Two tools cover MongoDB over MCP: For ad hoc reads of live data an agent uses query_run against connection.database.collection instead.

Permissions

MongoDB connections are an IAM resource type. Registering one requires Create on MongoDB connections → All MongoDB connections. Browsing a connection’s collections in the lake tree and importing from it require Describe on that specific connection, and scheduled syncs check the same permission for the principal that runs them. See Resources and actions.