Setup
- Open Settings → Lake and choose Add MongoDB connection.
- Paste a
mongodb://ormongodb+srv://connection string to fill the fields automatically, or enter them by hand:
The dialog includes the shell commands for a read-only user if you want to create one before saving.
Querying live
Collections are reachable asconnection_name.database.collection:
address.city is the column address_city. In the lake tree, field lists are sampled from the first 200 documents and _id is shown as the key.
A MongoDB collection selects DuckDB. DuckDB is the only engine that attaches external connections, so the query router routes any query naming a connection collection there before it considers input size - leave
engine on auto. Asking for Bloom, Polars, or Spark on one of these queries returns an engine capability error rather than rerouting silently.Importing into Iceberg
For anything larger than an ad hoc read, import the collection. The import runs as a Spark job using the MongoDB Spark connector and lands the data as an Iceberg table in your lake, where every engine can reach it. The schema is inferred from a sample of documents:- Nested documents are kept as structs and arrays as arrays, so the imported table is not flattened the way live queries are.
- A field seen with conflicting types is stored as a string. A field that is only ever null is also stored as a string.
- Integers and decimals are widened so later, larger values still fit.
_idis stored as a string, with ObjectIds as their 24-character hex.
oleander.default.<collection> - the sanitized collection name, with no connection prefix. OVERWRITE replaces it, APPEND adds to it.
An import does not return rows. It submits a job and returns a run_id you can monitor like any other run, and once it completes you can sample the destination with a normal query.
Collection syncing
A one-off import gives you a snapshot. Put the same import on a schedule and the Iceberg copy tracks the source instead. Schedules run hourly or daily, and are managed the same way as Postgres syncs: under Scheduled syncs in Settings → Lake, and under Transfers in the lake’s left panel.Scheduled syncs are incremental. After the first full import, each run merges new and changed documents on
_id rather than recopying the collection.- A timestamp field such as
updatedAt,updated_at,lastModifiedAt, ormodifiedAtmeans documents modified since the last run are merged. - Otherwise, if
_idvalues are ObjectIds, documents inserted since the newest one’s embedded creation time are merged. Updates to existing documents are not picked up on this path. - Anything else falls back to a full refresh.
From an agent
Two tools cover MongoDB over MCP:
For ad hoc reads of live data an agent uses
query_run against connection.database.collection instead.