Skip to main content
Lake Light

Overview

oleander’s lake provides a powerful environment for data analysis, collaboration, and discovery. Execute SQL queries, manage your tables, and contribute to the community with automatic lineage tracking and observability built-in. DuckDB is one of four engines behind the lake, and the most widely capable one. It is the only engine that reaches external connections and catalogs other than oleander, and it handles mutations that Bloom cannot commit. You do not select it by hand: the query router sends work here when the query calls for it. Every organization is automatically provisioned a private Apache Iceberg catalog named oleander. It has two namespaces: default for your own tables and telemetry for platform events (run events, traces, logs). It is cross-compatible with Spark, DuckDB, and any other Iceberg-aware tool. See Catalogs for the full schema. You can also connect BigQuery, Snowflake, and Postgres to query them alongside your lake data, bring your own compute to connect from a local DuckDB or PySpark session, or register an external Iceberg catalog to query your existing S3 Tables data. The lake’s left panel splits into two tabs: Data, for browsing your Iceberg catalogs and external connections, and Schedules, for scheduled queries, transfers (recurring Postgres syncs), and table maintenance, together with their cadence and health status. See Scheduled queries with time-travel below, Table syncing, and Table health and maintenance for what lives in each.

Portable

The lake is portable and can be used with any DuckDB environment. This means you can run queries, import and export data, and leverage all the features of oleander’s lake wherever DuckDB is supported, including on your local machine, in cloud environments, or as part of larger data workflows.
For full lineage coverage, always use the oleander lake API or platform (not a direct DuckDB connection) for queries. API usage ensures all operations are tracked and observable through the lake interface.Direct DuckDB connections are supported for power users; however, operations performed outside the API will not be traceable in lineage or workflows.

Automatic lineage

Every operation in oleander automatically captures lineage metadata using the OpenLineage specification. No manual configuration id required. Auto Lineage Light

Serverless SQL execution

Run SQL queries directly in oleander’s lake environment. Whether you’re analyzing your own data or exploring public tables, the SQL interface provides:
  • Fast Query Execution - Execute queries with optimized performance
  • Interactive Results - View and export query results instantly
  • Query History - Access your past queries and results
  • Automatic Lineage - Every query automatically captures lineage metadata

Sync your data

There are several ways to sync your own data into oleander’s lake:
Use the dialog to sync parquet files directly into your lakeUpload Dialog Light
Connect to your lake and run any DuckDB import operation to load data from various sources
Run CREATE TABLE statements to define and populate tables with your data
Use API-based queries to pull data from external sources directly into your lake (see example below)
Configure a daily sync of parquet files from your S3 bucket into oleander.default. Navigate to lake settings and add a sync under Scheduled S3 syncs. Provide your AWS credentials, bucket name, region, and optional prefix. Only files whose names don’t already exist as tables will be synced on each run.Lake Settings Light

Scheduled queries with time-travel

Any saved query can be put on a schedule. Open Schedule query on a query tab and set:
  • Runs - Every 15 minutes, Hourly, or Daily. The 15-minute cadence suits rollups whose work is bounded by what arrived since the last run, so a tighter interval is cheaper per run rather than more expensive.
  • Engine - Auto, Bloom, DuckDB, Spark, or Polars. Leave it on Auto unless you have a reason to pin one.
  • Destination - the namespace.table in the oleander catalog where each run lands, e.g. oleander.default.sch_nyc_taxi_daily.
Each successful run writes a new Iceberg snapshot to the destination table rather than overwriting it, so past runs stay queryable with time-travel instead of you needing to keep separate historical copies of the data. Please note that these do incur additional storage as each version is kept to preserve time-traveling capabilities.
Scheduled queries are listed under the lake panel’s Schedules tab, alongside transfers, with their cadence and next run time. A failed run shows what went wrong on hover. Schedules can also be created and managed from an agent - see Saved query schedules.

External connections

Attach BigQuery, Snowflake, and Postgres in lake settings and query them alongside your Iceberg data as connection_name.schema.table:
BigQuery, Snowflake, and Postgres always select DuckDB. DuckDB is the only engine that attaches external connections, so the query router sends any query naming a connection table here, whatever engine you asked for. Requesting Bloom, Polars, or Spark for one of these queries returns an engine capability error rather than rerouting silently.
Credentials are stored encrypted and are never returned to the browser after saving.

Bring your own (Iceberg) catalog

We support adding Iceberg via S3 table buckets. This lets you use your existing Iceberg catalog within the oleander lake and consume the oleander compute layer.
This allows the oleander aws account to generate temporary credentials to access your catalog.
Create a role that has access to S3Tables. You can scope this down to specific tables.
Configure the trust relationship for an your IAM role with the following policy:
Navigate to the lake settings and your aws account ID, role ARN, and bucket name. Next your catalog should be available listed on the lake page. Iceberg Dialog Light

Bring your own compute

Connect to your oleander lake from any DuckDB terminal, PySpark application, or client library. You can easily connect to your lake in our settings here.

DuckDB terminal

Use the oleander CLI to launch a connected DuckDB session:

Getting started

Get started with oleander’s lake in just a few steps:
  1. Navigate to the Lake in your oleander dashboard to access the SQL interface
  2. Run a SQL query on a public table to get familiar with the interface:
  1. Sync your own data to start analyzing your tables. Create a new table with sample data:
Follow up with an insertion if you’d like as well:
  1. Explore lineage to see how data flows through your queries. Every operation automatically captures lineage metadata for complete observability.
The lake provides a complete environment for data analysis, from quick SQL queries to complex Spark jobs, all with observability built in.

FAQs

Clock time and query time can differ for several reasons:
  • Clock time measures the total wall-clock time from when a query is submitted until it comes back to you. This also includes lineage extraction and creation.
  • Query time measures the actual execution time spent processing the query in the database engine running your query.
The difference typically occurs due to:
  • Network latency
  • Lineage extraction
You only pay for the query time.
Direct DuckDB connections bypass the oleander API layer, which is responsible for capturing and instrumenting lineage metadata. When you connect directly to DuckDB, your queries execute directly against the database engine without going through oleander’s observability infrastructure.To capture lineage:
  • Use the oleander lake API or platform for queries
  • All operations performed through the API are automatically tracked and observable
  • Lineage metadata is captured using the OpenLineage specification
Direct DuckDB connections are supported for power users who need direct database access, but they operate outside the observability layer. If you need lineage tracking for your queries, use the oleander API instead of a direct connection.