> ## Documentation Index
> Fetch the complete documentation index at: https://docs.oleander.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# MongoDB

> Query a live MongoDB database from the lake, and import collections into Iceberg once or on a schedule.

Register a MongoDB database and its collections become a live, read-only catalog in the lake. Query them directly for ad hoc work, or import a collection into Iceberg when you want a snapshot you can run large jobs against.

Connections are managed in [Settings → Lake](https://oleander.dev/app/settings/lake).

<Tip>
  Point oleander at a secondary or a dedicated read-only user, and make sure your cluster's network access list allows oleander's egress.
</Tip>

## Setup

1. Open [Settings → Lake](https://oleander.dev/app/settings/lake) and choose **Add MongoDB connection**.
2. Paste a `mongodb://` or `mongodb+srv://` connection string to fill the fields automatically, or enter them by hand:

| Field | Description |
| - | - |
| **Connection name** | Lowercase letters, numbers, and underscores, e.g. `my_mongo`. This becomes the catalog prefix in SQL. Names are shared across all connection types, so it cannot match an existing Postgres, MySQL, or other connection. |
| **Host** | Cluster or server host, e.g. `cluster0.abcde.mongodb.net` |
| **SRV record** | Check for `mongodb+srv://` hosts, which is what Atlas uses |
| **Port** | Defaults to `27017`. Not used with SRV |
| **Database** | The database to expose |
| **Options** | Extra URI options, e.g. `authSource=admin` |
| **Username** | The user oleander connects as |
| **Password** | Stored encrypted and never returned to the browser |

The dialog includes the shell commands for a read-only user if you want to create one before saving.

## Querying live

Collections are reachable as `connection_name.database.collection`:

```sql theme={null}
-- Read straight from MongoDB
SELECT _id, email, created_at
FROM my_mongo.app_production.users
WHERE created_at >= '2026-01-01'
LIMIT 100;

-- Join MongoDB against your Iceberg lake
SELECT u.email, count(*) AS events
FROM my_mongo.app_production.users u
JOIN oleander.default.events e ON e.user_id = u._id
GROUP BY 1;
```

Nested documents are flattened for SQL: a field at `address.city` is the column `address_city`. In the lake tree, field lists are sampled from the first 200 documents and `_id` is shown as the key.

<Note>
  **A MongoDB collection selects DuckDB.** DuckDB is the only engine that attaches external connections, so the [query router](/platform/query-routing/overview) routes any query naming a connection collection there before it considers input size - leave `engine` on `auto`. Asking for Bloom, Polars, or Spark on one of these queries returns an engine capability error rather than rerouting silently.
</Note>

Live reads run against your database, so filter and limit the way you would against production. Once a collection is [imported](#importing-into-iceberg) it lives in Iceberg and the router is free to send it to any engine.

## Importing into Iceberg

For anything larger than an ad hoc read, import the collection. The import runs as a Spark job using the MongoDB Spark connector and lands the data as an Iceberg table in your lake, where every engine can reach it.

The schema is inferred from a sample of documents:

* Nested documents are kept as structs and arrays as arrays, so the imported table is not flattened the way live queries are.
* A field seen with conflicting types is stored as a string. A field that is only ever null is also stored as a string.
* Integers and decimals are widened so later, larger values still fit.
* `_id` is stored as a string, with ObjectIds as their 24-character hex.

The destination defaults to `oleander.default.<collection>` - the sanitized collection name, with no connection prefix. `OVERWRITE` replaces it, `APPEND` adds to it.

An import does not return rows. It submits a job and returns a `run_id` you can monitor like any other run, and once it completes you can sample the destination with a normal query.

## Collection syncing

A one-off import gives you a snapshot. Put the same import on a schedule and the Iceberg copy tracks the source instead. Schedules run **hourly** or **daily**, and are managed the same way as [Postgres syncs](/connections/postgres#table-syncing): under **Scheduled syncs** in [Settings → Lake](https://oleander.dev/app/settings/lake), and under **Transfers** in the lake's left panel.

<Note>
  **Scheduled syncs are incremental.** After the first full import, each run merges **new and changed documents** on `_id` rather than recopying the collection.
</Note>

How a run finds what changed depends on the collection, and the reason is written to the run's plan log:

* A timestamp field such as `updatedAt`, `updated_at`, `lastModifiedAt`, or `modifiedAt` means documents modified since the last run are merged.
* Otherwise, if `_id` values are ObjectIds, documents inserted since the newest one's embedded creation time are merged. Updates to existing documents are not picked up on this path.
* Anything else falls back to a full refresh.

New top-level fields that appear in the source are added as columns; new fields inside an existing struct wait for the next full refresh. **Deletes are never synced.** A document removed from MongoDB stays in the Iceberg copy.

Every sync run is a normal oleander run: it appears in run history, emits lineage with the collection as the input and the Iceberg table as the output, and carries cost like any other job.

## From an agent

Two tools cover MongoDB over [MCP](/mcp/introduction):

| Tool | What it does |
| - | - |
| `mongo_connections_list` | List registered connections - name, host, port, SRV, database, username, options. Credentials are never returned. |
| `mongo_collections_import` | Import a collection into Iceberg with the sampled-schema Spark read described above |

For ad hoc reads of live data an agent uses `query_run` against `connection.database.collection` instead.

## Permissions

MongoDB connections are an IAM resource type. Registering one requires **Create** on **MongoDB connections → All MongoDB connections**. Browsing a connection's collections in the lake tree and importing from it require **Describe** on that specific connection, and scheduled syncs check the same permission for the principal that runs them. See [Resources and actions](/iam/resources-and-actions#connections).
