Connecting query engines to AIStor Tables

AIStor Tables exposes a native Iceberg REST Catalog so query and analytics engines can read and write Iceberg tables directly in MinIO AIStor object storage. This page consolidates connection details for the most common engines, and provides complete, working catalog configurations.

The catalog is served from the /_iceberg base path on the MinIO AIStor S3 endpoint. For example, a cluster reachable at https://aistor.example.net:9000 serves its catalog at https://aistor.example.net:9000/_iceberg.

How authentication works

The AIStor Tables catalog accepts two forms of authentication.

AWS Signature Version 4 (SigV4) request signing with the signing (service) name s3tables. Most Iceberg clients and query engines use this form.

OAuth2 bearer tokens, for engines that cannot sign catalog requests with SigV4. A client obtains a token from the catalog’s own token endpoint, or presents a token issued by the configured OpenID provider:

  • The catalog exposes an OAuth2 client-credentials token endpoint at POST /_iceberg/v1/oauth/tokens. The client_id and client_secret are an AIStor access key and secret key, which may belong to the root user, an IAM user, or a service account. The endpoint returns an AIStor session token that the client then sends as Authorization: Bearer.
  • A bearer token issued by the configured OpenID provider is also accepted. The catalog validates it under the same rules the Security Token Service applies and runs the request as the provider identity, with the policy the token’s claim names. An engine that forwards the end user’s provider token, such as StarRocks, Trino, or the Apache Iceberg client, therefore needs no separate exchange at AssumeRoleWithWebIdentity.

Bearer authentication and the token endpoint are enabled by default. See Iceberg REST catalog settings to turn them off or change the token validity.

The region value that an engine sends with SigV4 is required by the signing algorithm but is not otherwise used by AIStor Tables. Any non-empty value (for example, local or dummy) is acceptable.

For details on the underlying REST API and SigV4 headers, see the AIStor Tables API Reference. For policy-based access control of catalog actions, see Controlling access to AIStor Tables.

Supported engines

Engine Catalog authentication Notes
Apache Spark (Iceberg REST) SigV4 (s3tables) Use the Iceberg REST catalog with rest.sigv4-enabled and rest.signing-name=s3tables.
PyIceberg SigV4 (s3tables) Use rest.sigv4-enabled and rest.signing-name=s3tables.
Trino / Presto SigV4 (s3tables) Property name for SigV4 differs by Trino version. See Trino.
Starburst SigV4 (s3tables) Built on Trino. Use the Trino Iceberg REST connector with SigV4.
Dremio SigV4 (s3tables) Requires Dremio 26.0 or later for the Iceberg REST Catalog source. See Dremio.
DuckDB SigV4 (s3tables) Requires DuckDB 1.5 or later. Nested namespaces are not supported. See DuckDB.
ClickHouse OAuth2 bearer Cannot sign catalog requests with SigV4. Authenticate it with a bearer token from the catalog’s token endpoint, which is enabled by default.
PuppyGraph SigV4 (s3tables) Connect through the Iceberg REST catalog with SigV4 signing.

Shared configuration values

All examples use the following placeholder values. Replace them with values for your deployment:

Value Description
uri / catalog URI The MinIO AIStor S3 endpoint with the /_iceberg catalog path, for example https://aistor.example.net:9000/_iceberg.
warehouse The plain warehouse name (for example analytics). Do not prefix it with s3:// or s3a://.
s3.endpoint The MinIO AIStor S3 endpoint, for example https://aistor.example.net:9000.
region Required by SigV4 but unused by AIStor. Use any non-empty value such as local.
access key / secret key Credentials for a user with permission to access AIStor Tables.
Path-style access
Always enable path-style S3 access (s3.path-style-access=true and the equivalent Hadoop S3A setting). Virtual-host-style addressing is not used for warehouse buckets in these configurations.

PyIceberg

PyIceberg connects to the REST catalog with SigV4 signing. Install the dependencies with:

pip install pyiceberg pyarrow pandas

Load the catalog:

from pyiceberg.catalog import load_catalog

catalog = load_catalog(
    "aistor",
    **{
        "uri": "https://aistor.example.net:9000/_iceberg",
        "warehouse": "analytics",
        "rest.sigv4-enabled": "true",
        "rest.signing-name": "s3tables",
        "rest.signing-region": "local",   # required by SigV4, value unused
        "client.region": "local",
        "client.access-key-id": "YOUR-ACCESS-KEY",
        "client.secret-access-key": "YOUR-SECRET-KEY",
        "s3.endpoint": "https://aistor.example.net:9000",
        "s3.path-style-access": "true",
        "s3.access-key-id": "YOUR-ACCESS-KEY",
        "s3.secret-access-key": "YOUR-SECRET-KEY",
    }
)

To authenticate with a bearer token instead, give PyIceberg a credential and the catalog’s token endpoint:

from pyiceberg.catalog.rest import RestCatalog

catalog = RestCatalog(
    "aistor",
    **{
        "uri": "https://aistor.example.net:9000/_iceberg",
        "warehouse": "analytics",
        "credential": "YOUR-ACCESS-KEY:YOUR-SECRET-KEY",
        "oauth2-server-uri": "https://aistor.example.net:9000/_iceberg/v1/oauth/tokens",
        # Data-file reads still need S3 credentials unless the server vends them.
        "s3.endpoint": "https://aistor.example.net:9000",
        "s3.path-style-access": "true",
        "s3.access-key-id": "YOUR-ACCESS-KEY",
        "s3.secret-access-key": "YOUR-SECRET-KEY",
    }
)

For a complete end-to-end PyIceberg walkthrough that creates a warehouse, namespace, and table and then inserts and queries data, see AIStor Tables.

Spark

Spark uses the Iceberg Spark runtime with the REST catalog and SigV4 signing. The example below configures a catalog named aistor.

config = {
    # Catalog definition
    "spark.sql.catalog.aistor": "org.apache.iceberg.spark.SparkCatalog",
    "spark.sql.catalog.aistor.type": "rest",
    "spark.sql.catalog.aistor.uri": "https://aistor.example.net:9000/_iceberg",
    "spark.sql.catalog.aistor.warehouse": "analytics",

    # REST catalog SigV4 signing
    "spark.sql.catalog.aistor.rest.endpoint": "https://aistor.example.net:9000",
    "spark.sql.catalog.aistor.rest.access-key-id": "YOUR-ACCESS-KEY",
    "spark.sql.catalog.aistor.rest.secret-access-key": "YOUR-SECRET-KEY",
    "spark.sql.catalog.aistor.rest.sigv4-enabled": "true",
    "spark.sql.catalog.aistor.rest.signing-name": "s3tables",
    "spark.sql.catalog.aistor.rest.signing-region": "local",  # required, value unused

    # S3 data access
    "spark.sql.catalog.aistor.s3.endpoint": "https://aistor.example.net:9000",
    "spark.sql.catalog.aistor.s3.access-key-id": "YOUR-ACCESS-KEY",
    "spark.sql.catalog.aistor.s3.secret-access-key": "YOUR-SECRET-KEY",
    "spark.sql.catalog.aistor.s3.path-style-access": "true",
    "spark.sql.catalog.aistor.io-impl": "org.apache.iceberg.aws.s3.S3FileIO",

    # Iceberg extensions and runtime JARs
    "spark.sql.extensions": "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions",
    "spark.sql.defaultCatalog": "aistor",
    "spark.jars.packages": (
        "org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.10.1,"
        "org.apache.iceberg:iceberg-aws-bundle:1.10.1"
    ),
}

Match the iceberg-spark-runtime artifact to your Spark and Scala versions (for example, iceberg-spark-runtime-3.5_2.12 for Spark 3.5 with Scala 2.12). The iceberg-aws-bundle artifact provides the AWS SDK and S3 FileIO that Spark uses for data access.

Trino

Trino connects to the catalog through its Iceberg connector with SigV4 signing. The property keys are the same in both formats; only the SigV4 enablement property differs by Trino version.

SigV4 property differs by Trino version
  • Trino 477 and later: use iceberg.rest-catalog.security=SIGV4.
  • Trino 476 and earlier: use iceberg.rest-catalog.sigv4-enabled=true.

Set the form that matches your Trino version. Do not set both.

Static catalog properties file

Place the following in an iceberg.properties file in the Trino catalog directory (typically etc/catalog/iceberg.properties). This example targets Trino 477 or later.

connector.name=iceberg
iceberg.catalog.type=rest
iceberg.rest-catalog.uri=https://aistor.example.net:9000/_iceberg
iceberg.rest-catalog.warehouse=analytics
iceberg.rest-catalog.security=SIGV4
iceberg.rest-catalog.signing-name=s3tables
iceberg.rest-catalog.vended-credentials-enabled=true
iceberg.rest-catalog.view-endpoints-enabled=true
iceberg.unique-table-location=true
s3.region=local
s3.endpoint=https://aistor.example.net:9000
s3.aws-access-key=YOUR-ACCESS-KEY
s3.aws-secret-key=YOUR-SECRET-KEY
s3.path-style-access=true
fs.hadoop.enabled=false
fs.native-s3.enabled=true

For Trino 476 or earlier, replace the SigV4 line:

iceberg.rest-catalog.sigv4-enabled=true

Dynamic catalog creation (SQL)

If your Trino deployment has the CREATE CATALOG SQL syntax enabled, you can create the catalog at runtime. This example targets Trino 477 or later.

CREATE CATALOG aistor USING iceberg
WITH (
    "iceberg.catalog.type" = 'rest',
    "iceberg.rest-catalog.uri" = 'https://aistor.example.net:9000/_iceberg',
    "iceberg.rest-catalog.warehouse" = 'analytics',
    "iceberg.rest-catalog.security" = 'SIGV4',
    "iceberg.rest-catalog.signing-name" = 's3tables',
    "iceberg.rest-catalog.vended-credentials-enabled" = 'true',
    "iceberg.rest-catalog.view-endpoints-enabled" = 'true',
    "iceberg.unique-table-location" = 'true',
    "s3.region" = 'local',
    "s3.endpoint" = 'https://aistor.example.net:9000',
    "s3.aws-access-key" = 'YOUR-ACCESS-KEY',
    "s3.aws-secret-key" = 'YOUR-SECRET-KEY',
    "s3.path-style-access" = 'true',
    "fs.hadoop.enabled" = 'false',
    "fs.native-s3.enabled" = 'true'
);

For Trino 476 or earlier, replace the "iceberg.rest-catalog.security" = 'SIGV4' line with "iceberg.rest-catalog.sigv4-enabled" = 'true'.

TLS and the Java truststore

When MinIO AIStor serves the catalog over HTTPS with a certificate signed by an internal or self-signed certificate authority (CA), the Java runtime that Trino uses must trust that CA. Otherwise, Trino fails catalog connections with a PKIX or “unable to find valid certification path” error.

Import the CA certificate into the truststore that the Trino JVM uses, for example:

keytool -import -alias aistor-ca \
  -file aistor-ca.crt \
  -keystore "$JAVA_HOME/lib/security/cacerts" \
  -storepass changeit

Restart Trino after updating the truststore. Certificates issued by a well-known public CA are already trusted by the default Java truststore and do not require this step.

Starburst

Starburst is built on Trino and uses the same Iceberg REST connector and properties shown in the Trino section, including the version-specific SigV4 property and the same Java truststore requirement for internal CAs.

Dremio

Dremio connects through its Iceberg REST Catalog source with SigV4 signing and the s3tables signing name.

Dremio 26.0 or later is required
The generic Iceberg REST Catalog source type was added in Dremio 26.0. The source type is also absent from the Dremio community build (dremio-oss).

Add a new source of type Iceberg REST Catalog and set the following.

General

Field Value
Name Any label, for example aistor
Endpoint URI https://aistor.example.net:9000/_iceberg
Use vended credentials Cleared

Advanced Options - Catalog Properties

warehouse                       = analytics
rest.sigv4-enabled              = true
rest.signing-name               = s3tables
rest.signing-region             = local
dremio.s3.region                = local
fs.s3a.endpoint                 = https://aistor.example.net:9000
fs.s3a.path.style.access        = true
fs.s3a.aws.credentials.provider = org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider

Advanced Options - Catalog Credentials

rest.access-key-id     = YOUR-ACCESS-KEY
rest.secret-access-key = YOUR-SECRET-KEY
fs.s3a.access.key      = YOUR-ACCESS-KEY
fs.s3a.secret.key      = YOUR-SECRET-KEY

As with Trino, a certificate signed by an internal CA must be trusted by the Java runtime Dremio uses. See TLS and the Java truststore.

DuckDB

DuckDB connects to the catalog through its iceberg extension with SigV4 signing. Use DuckDB 1.5 or later: earlier releases cannot set the SigV4 signing name and cannot connect.

INSTALL iceberg; LOAD iceberg;
INSTALL httpfs; LOAD httpfs;

CREATE SECRET aistor_s3 (
    TYPE S3,
    KEY_ID 'YOUR-ACCESS-KEY',
    SECRET 'YOUR-SECRET-KEY',
    ENDPOINT 'aistor.example.net:9000',   -- host and port only, no scheme
    URL_STYLE 'path',
    USE_SSL 1,
    REGION 'local'                        -- required, value unused
);

ATTACH 'analytics' AS aistor (
    TYPE ICEBERG,
    ENDPOINT 'https://aistor.example.net:9000/_iceberg',
    AUTHORIZATION_TYPE 'sigv4',
    SIGV4_SERVICE 's3tables',
    SIGV4_REGION 'local',
    SECRET aistor_s3
);

Always set SIGV4_SERVICE. Without it, DuckDB tries to read the signing name from an AWS-style hostname and fails with Could not parse AWS service from host. Do not use ENDPOINT_TYPE 's3_tables', which builds an s3tables.<region>.amazonaws.com endpoint instead of using your deployment endpoint.

To authenticate with a bearer token instead, put the catalog endpoint on an ICEBERG secret. DuckDB resolves the token when the secret is created, not when the catalog is attached:

CREATE SECRET aistor_catalog (
    TYPE ICEBERG,
    CLIENT_ID 'YOUR-ACCESS-KEY',
    CLIENT_SECRET 'YOUR-SECRET-KEY',
    ENDPOINT 'https://aistor.example.net:9000/_iceberg'
);

ATTACH 'analytics' AS aistor (
    TYPE ICEBERG,
    SECRET aistor_catalog,
    ACCESS_DELEGATION_MODE 'none'
);

ACCESS_DELEGATION_MODE defaults to vended_credentials, so set it to none unless the server vends credentials. Vending is off by default; see Iceberg REST catalog settings. With vending enabled, omit both ACCESS_DELEGATION_MODE and the S3 secret, because the catalog returns the storage credentials itself.

Nested namespaces are not supported

DuckDB cannot run table operations in a namespace that has more than one level, such as finance.reporting. Those requests fail with 403 and the message The request signature we calculated does not match the signature you provided. DuckDB signs the namespace separator in the request path differently than the catalog does. Single-level namespaces are not affected, and other engines on this page are not affected.

Keep namespaces single-level for warehouses that DuckDB queries. To read an existing table in a nested namespace, scan it by location instead of through the catalog:

SET unsafe_enable_version_guessing = true;
SELECT * FROM iceberg_scan('s3://analytics/<table-location>');

Find the location with mc table info. This path is read-only and bypasses the catalog.

ClickHouse

ClickHouse connects through its DataLakeCatalog database engine. It cannot sign catalog requests with SigV4, so it authenticates with an OAuth2 bearer token that it obtains from the catalog’s token endpoint.

SET allow_database_iceberg = 1;

CREATE DATABASE aistor
ENGINE = DataLakeCatalog('https://aistor.example.net:9000/_iceberg')
SETTINGS
    catalog_type = 'rest',
    warehouse = 'analytics',
    catalog_credential = 'YOUR-ACCESS-KEY:YOUR-SECRET-KEY',
    oauth_server_uri = 'https://aistor.example.net:9000/_iceberg/v1/oauth/tokens',
    storage_endpoint = 'https://aistor.example.net:9000';

SHOW TABLES IN aistor;

catalog_credential takes the access key and secret key as KEY:SECRET. ClickHouse exchanges them at oauth_server_uri for a token, and sends that token as Authorization: Bearer on every catalog request.

The token endpoint is enabled by default. See Iceberg REST catalog settings to confirm it is on, or to change how long an issued token stays valid.

storage_endpoint is the S3 endpoint ClickHouse uses to read the data files. If you turn on vended credentials, add vended_credentials = true so that ClickHouse takes the storage credentials from the catalog response instead of its own configuration.