Lake tables: Iceberg, Delta and Parquet
How to let people query Apache Iceberg, Delta Lake and Parquet tables in your Amazon S3 storage with read-only SQL, when you have no query engine (such as Athena, Trino, Snowflake or Databricks) for SourceLace to ask. SourceLace reads the table files itself, so an admin must turn this on, and only for tables that have no row or column rules.
Status: Preview. New, offered for pilots and provided as is. Built from the Iceberg and Delta specifications and AWS's documentation, and covered by automated tests with real table files in a simulated S3 bucket; not yet used with a customer's live account.
At a glance
| Connects through | Sign-in | Network | Status | First test question | |
|---|---|---|---|---|---|
| Lake tables | Amazon S3 API, with AWS Glue Data Catalog or an Iceberg REST catalog | Shared account | Nothing to open | Preview | "Describe the orders table in lake:warehouse." |
Add each source with the usual steps. The network answers are explained in What your network needs. If the first test question fails, the error message tells you what to fix: look it up under When something goes wrong below.
When to use it
If your tables already have a query engine, connect that instead (for example Snowflake or Databricks): the engine applies each person's own permissions, including row and column rules, and does the heavy work.
Use a lake source when there is no engine, for tables everyone allowed to use the source may see in full. It reads the files directly, which skips any row filters, column masks or per-person grants that AWS Lake Formation, Unity Catalog, Ranger or a query engine would apply. That is why it only works once an admin sets the direct_reads option to on.
Like Amazon S3, a lake source has no per-person sign-in: everyone who may use it reads with your organization's own AWS access, set on the source. To narrow who in your organization may use it, use access by group.
What people can do
search_schemalists the tables the admin declared;describe_objectlists a table's columns and types, from the table's own metadata.query(languagesql) takes oneSELECTorWITH ... SELECTstatement in DuckDB's SQL dialect over the declared tables, by name, such asSELECT region, sum(amount) FROM orders WHERE order_date >= DATE '2026-01-01' GROUP BY region. Joins across the source's tables work, whatever their format.- Anything else is refused before any file is read: changes, more than one statement, other tables, file paths or web addresses, and functions that read files or settings (
read_parquet,read_csv,glob,queryand the like). - SourceLace reads only the columns a query uses and skips the parts of files, and whole files, that a filter rules out (for example a filter on a partition column), so name the columns you need and filter on partition columns.
- Results are capped like other queries (2,000 rows kept, 50 shown at a time), and held for at most 30 minutes and never written to a database (see How long things are kept).
- Lake sources never accept changes.
Supported tables:
| Format | How to name it | Notes |
|---|---|---|
| Iceberg, by its folder or metadata file | orders = iceberg sales/orders/ |
The newest metadata file is found from version-hint.text, or the highest vN.metadata.json. Renamed and added columns are followed. |
| Iceberg, in AWS Glue Data Catalog | orders = iceberg glue:sales_db.orders |
SourceLace asks Glue (GetTable) for the table's current metadata file. |
| Iceberg, in an Iceberg REST catalog | orders = iceberg rest:sales.orders |
Needs the catalog_uri option, and usually catalog_token. |
| Delta Lake | customers = delta crm/customers/ |
Read from the latest checkpoint and the commits after it. Column mapping is supported. |
| Parquet | clicks = parquet raw/clicks/ |
Every .parquet file in the folder (or one file). key=value folders become extra text columns. The files must share one schema. |
Not supported yet: Iceberg tables with row-level deletes (delete files or deletion vectors), Delta tables with deletion vectors or V2 checkpoints, ORC and Avro data files, time travel, and storage other than Amazon S3. Such a table is refused with a message saying so.
1. Give SourceLace read access in AWS (AWS admin)
SourceLace needs to list and read the objects under the tables' location, and, for tables in Glue, to read their definitions. Create an IAM policy like this one, with your bucket and folder:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::corvanta-lake",
"Condition": { "StringLike": { "s3:prefix": ["warehouse/*"] } }
},
{
"Effect": "Allow",
"Action": "s3:GetObject",
"Resource": "arn:aws:s3:::corvanta-lake/warehouse/*"
},
{
"Effect": "Allow",
"Action": "glue:GetTable",
"Resource": [
"arn:aws:glue:us-east-1:123456789012:catalog",
"arn:aws:glue:us-east-1:123456789012:database/sales_db",
"arn:aws:glue:us-east-1:123456789012:table/sales_db/*"
]
}
]
}
Leave out the Glue statement if no table is in Glue. If the bucket uses a customer-managed KMS key, also allow kms:Decrypt on that key.
Then give SourceLace the policy in one of two ways, exactly as for an Amazon S3 source:
- A role (recommended). Create an IAM role with the policy, trusting SourceLace's AWS account, with a trust policy condition that requires the external ID SourceLace shows for your organization on the source (once saved). SourceLace assumes the role for each query and never stores the short-lived credentials.
- Access keys. An IAM user with the policy and an access key. The secret is stored encrypted and never shown again.
2. Add the source (SourceLace admin)
On Manage sources, add a source of kind Lake tables (lake), such as lake:warehouse, with:
| Option | What to enter |
|---|---|
location |
Where the tables are, such as s3://corvanta-lake/warehouse/. SourceLace never reads outside it, even if a table's metadata points elsewhere. |
region |
The bucket's AWS region, such as us-east-1. |
tables |
The tables, separated by semicolons (or new lines), each name = format where, as in the table above. Folders are relative to the location, or a full s3:// address inside it. Table names are lowercase letters, digits and underscores. |
role_arn, or access_key_id and secret_access_key |
Your organization's AWS access from step 1. |
direct_reads |
on: you confirm these tables have no row or column rules. Without it, the source is refused. |
catalog_uri |
For rest: tables: the catalog's https address, such as https://catalog.corvanta.com/api/catalog. |
catalog_warehouse |
For rest: tables, if your catalog needs one. |
catalog_token |
For rest: tables: a bearer token with read access to the tables' metadata. Stored encrypted, never shown again, and dropped if catalog_uri changes. |
max_bytes_read, max_files_read |
Optional: lower this source's caps (below). |
Then connect it once like any other source (there is no sign-in page) and try describe_object on a table.
Limits and safety
- Bytes and files. One query may read at most 2 GB and 2,000 files (data and metadata together) unless your limits say otherwise; a source's
max_bytes_readandmax_files_readoptions can only lower them. SourceLace counts every read as it happens and stops the query at the cap, so the cap is never exceeded. - Time. Queries have the usual query time limit and are stopped when it runs out.
- Kept apart. Each query runs on its own and shares nothing with other queries, people or organizations. It reads only the files of the tables you declared, inside the location, and cannot reach the internet or other buckets.
- Working space. A very large query may use temporary working space while it runs. That space is not encrypted and is deleted when the query ends. Nothing else is written down.
When something goes wrong
- "... skips any row or column rules ...": set
direct_readstoon, but only if the tables have no such rules. - "Amazon S3 refused ... access": the role or keys cannot read that table's files. Check the IAM policy (and KMS key) covers the whole location.
- "... was not found": check the table's entry in
tables(folder, Glue database and table, or catalog namespace). - "... names a file outside the source's location": the table's metadata points outside
location. Widenlocationif that is expected. - "... would read more than ...": select fewer columns, filter on partition columns, or ask an admin to raise the limit.
- "... cannot read directly yet": the table uses a feature listed above as not supported. Query it through your query engine.
- "The Iceberg catalog refused ...": enter a new
catalog_token.