Tables and Delta Sharing Observability
AIStor exposes metrics, audit-log entries, and alerting hooks for the AIStor Tables (Apache Iceberg REST Catalog) and AIStor Table Sharing (Delta Sharing) APIs. Use these signals to monitor request volume, latency, authentication, caching, and rate limiting for table and share workloads, and to monitor the catalog scanner on a disaster-recovery replica.
Metrics
Tables and Delta Sharing publish metrics on both the v3 and v2 metrics endpoints.
Version 3 metrics
The v3 endpoints provide the full catalog of request, error, authentication, caching, and rate-limiting metrics for these APIs.
- Delta Sharing metrics are available under
/minio/metrics/v3/delta-sharing. See Delta Sharing metrics for the complete reference. - Tables metrics are available under
/minio/metrics/v3/tables. See Tables metrics for the complete reference.
A scrape job using the root /minio/metrics/v3 endpoint captures both sets.
The following v3 metrics are the primary signals for monitoring these APIs:
| Name | Type | Description | Labels |
|---|---|---|---|
minio_delta_sharing_total |
counter | Total number of Delta Sharing API requests processed. | name, type, server |
minio_delta_sharing_errors_total |
counter | Total number of Delta Sharing API requests that resulted in errors. | name, type, server |
minio_delta_sharing_4xx_errors_total |
counter | Total number of Delta Sharing API requests that resulted in 4xx errors. | name, type, server |
minio_delta_sharing_5xx_errors_total |
counter | Total number of Delta Sharing API requests that resulted in 5xx errors. | name, type, server |
minio_delta_sharing_canceled_total |
counter | Total number of Delta Sharing API requests canceled by the client. | name, type, server |
minio_delta_sharing_inflight_total |
gauge | Current number of Delta Sharing API requests actively being processed. | name, type, server |
minio_delta_sharing_auth_success_total |
counter | Total successful Delta Sharing authentications. | server |
minio_delta_sharing_auth_failures_total |
counter | Total failed Delta Sharing authentications. | server |
minio_delta_sharing_oauth_tokens_issued_total |
counter | Total OAuth tokens issued for Delta Sharing. | server |
minio_delta_sharing_cache_hits_total |
counter | Total Delta Sharing cache hits. | cache_type, server |
minio_delta_sharing_cache_misses_total |
counter | Total Delta Sharing cache misses. | cache_type, server |
minio_delta_sharing_cache_size |
gauge | Current size of the Delta Sharing cache. | cache_type, server |
minio_delta_sharing_rate_limited_total |
counter | Total number of rate-limited Delta Sharing requests. | server |
minio_delta_sharing_requests_ttfb_seconds_distribution |
counter | Histogram distribution of time to first byte for Delta Sharing requests. | api, le, server |
minio_tables_total |
counter | Total number of Tables API requests processed. | name, type, server |
minio_tables_5xx_errors_total |
counter | Total number of Tables API requests that resulted in 5xx errors. | name, type, server |
minio_tables_canceled_total |
counter | Total number of Tables API requests canceled by the client. | name, type, server |
minio_tables_inflight_total |
gauge | Current number of Tables API requests actively being processed. | name, type, server |
minio_tables_requests_ttfb_seconds_distribution |
counter | Histogram distribution of time to first byte for Tables requests. | api, le, server |
For the per-operation counters and gauges, the name label carries the operation name and the type label is delta-sharing (for Delta Sharing) or the Tables API type.
For the cache metrics, the cache_type label is either token or snapshot.
Replica catalog scanner
On a disaster-recovery replica, the catalog scanner rebuilds the Iceberg catalog from replicated data. See AIStor Tables site replication and disaster recovery.
The scanner publishes the following metrics under /minio/metrics/v3/tables:
| Name | Type | Description | Labels |
|---|---|---|---|
minio_tables_catalog_scanner_cycle_total |
counter | Total completed catalog scanner cycles. | server |
minio_tables_catalog_scanner_running |
gauge | Whether a cycle is running now (0 or 1). | server |
minio_tables_catalog_scanner_last_cycle_duration_seconds |
gauge | Duration of the last completed cycle, in seconds. | server |
minio_tables_catalog_scanner_last_run_timestamp |
gauge | Unix timestamp of the last completed cycle. | server |
minio_tables_catalog_scanner_cycle_errors_total |
counter | Total cycles that failed with a non-cancellation error. | server |
minio_tables_catalog_scanner_actions_total |
counter | Total catalog mutations the scanner applied. The action label is created, updated, or tombstoned. |
action, server |
minio_tables_catalog_scanner_warehouses_last_cycle |
gauge | Warehouse buckets scanned in the last cycle. | server |
Only one node per cluster runs the scanner, and only that node publishes non-zero values.
Aggregate across nodes rather than alerting per server, and expect the reporting node to change when leadership moves.
A primary site does not run the scanner, so these metrics stay at zero there.
These metrics report whether the scanner is healthy.
They do not report how far behind any individual table is — for that, run mc table replicate status against the replica.
Version 2 metrics
The v2 endpoints expose latency histograms for both APIs.
These metrics do not use the minio_ prefix and carry only the api label, which holds the operation name (for example, LoadTable or QueryTable).
| Name | Type | Description | Labels |
|---|---|---|---|
tables_ttfb_seconds |
histogram | Time taken by the Tables/Iceberg API from request read to response write. | api |
delta_sharing_ttfb_seconds |
histogram | Time taken by the Delta Sharing API from request read to response write. | api |
As standard Prometheus histograms, each exposes _bucket, _sum, and _count series.
See Metrics version 2 for v2 scrape endpoints.
Audit logs
When audit logging is configured, AIStor records an audit event for each Tables and Delta Sharing API call.
- Tables (Iceberg REST) operations are recorded under the
tablessubsystem. The operation name identifies the action, for exampleCreateWarehouse,CreateTable,LoadTable, orQueryTable. - Delta Sharing API calls are recorded for each operation, such as
ListShares,GetShare,QueryTable, orOAuthToken. Each Delta Sharing audit entry includes the share token identifier in atokenfield, which you can use to attribute activity to a specific share recipient.
AIStor does not publish audit logs to any destination by default. Configure a Kafka or webhook target to receive these events.
Alerting
The following examples use Prometheus AlertManager rule formatting, consistent with the other alerts in this documentation.
Most of the metrics below are counters, and those examples use rate() over a 5-minute window.
The replica catalog scanner also publishes gauges: minio_tables_catalog_scanner_last_run_timestamp holds a Unix timestamp, so its alert compares it against time() instead of taking a rate.
Tune the thresholds and durations to match your workload baseline.
Delta Sharing authentication failures
This alert triggers if Delta Sharing authentication failures increase beyond the configured threshold.
A sustained increase can indicate misconfigured recipients, expired share credentials, or unauthorized access attempts.
The auth_failures_total metric carries only the server label, so the rule does not filter by operation.
alert: DeltaSharingAuthFailures
expr: rate(minio_delta_sharing_auth_failures_total[5m]) > 1
for: 2m
labels:
severity: warning
annotations:
summary: "Delta Sharing auth failures on {{ $labels.server }}: {{ $value | humanize }}/sec"
impact: "Recipients may be unable to access shares, or unauthorized access is being attempted."
action: "Check share credentials and recipient configuration. Review audit logs for the failing tokens."
This alert requires the Prometheus scraping configuration capture the metrics provided in the following API endpoint(s):
/minio/metrics/v3/delta-sharing
A scrape job using the root /minio/metrics/v3 endpoint satisfies the above requirement.
Delta Sharing server error rate
This alert triggers if the rate of 5xx errors on the Delta Sharing API increases beyond the configured threshold.
5xx errors indicate that AIStor failed to process a Delta Sharing request.
The 5xx_errors_total metric carries the name and type labels, so you can attribute errors to a specific operation.
alert: DeltaSharingServerErrorRate
expr: rate(minio_delta_sharing_5xx_errors_total[5m]) > 1
for: 2m
labels:
severity: critical
annotations:
summary: "Delta Sharing 5xx errors for {{ $labels.name }} on {{ $labels.server }}: {{ $value | humanize }}/sec"
impact: "Server-side failures cause share queries to fail for recipients."
action: "Check {{ $labels.server }} logs. Investigate backend storage or resource exhaustion."
This alert requires the Prometheus scraping configuration capture the metrics provided in the following API endpoint(s):
/minio/metrics/v3/delta-sharing
A scrape job using the root /minio/metrics/v3 endpoint satisfies the above requirement.
Delta Sharing requests rate limited
This alert triggers if Delta Sharing requests are being rate limited.
A sustained increase indicates that recipients are exceeding configured request limits, which may degrade their experience.
The rate_limited_total metric carries only the server label.
alert: DeltaSharingRateLimited
expr: rate(minio_delta_sharing_rate_limited_total[5m]) > 1
for: 5m
labels:
severity: warning
annotations:
summary: "Delta Sharing requests rate limited on {{ $labels.server }}: {{ $value | humanize }}/sec"
impact: "Recipients are being throttled and may experience failed or delayed share queries."
action: "Review rate-limit configuration and recipient request patterns."
This alert requires the Prometheus scraping configuration capture the metrics provided in the following API endpoint(s):
/minio/metrics/v3/delta-sharing
A scrape job using the root /minio/metrics/v3 endpoint satisfies the above requirement.
Tables server error rate
This alert triggers if the rate of 5xx errors on the Tables (Iceberg REST) API increases beyond the configured threshold.
The 5xx_errors_total metric carries the name and type labels, so you can attribute errors to a specific operation such as CreateTable or LoadTable.
alert: TablesServerErrorRate
expr: rate(minio_tables_5xx_errors_total[5m]) > 1
for: 2m
labels:
severity: critical
annotations:
summary: "Tables 5xx errors for {{ $labels.name }} on {{ $labels.server }}: {{ $value | humanize }}/sec"
impact: "Server-side failures cause catalog operations to fail for query engines."
action: "Check {{ $labels.server }} logs. Investigate backend storage or transaction recovery."
This alert requires the Prometheus scraping configuration capture the metrics provided in the following API endpoint(s):
/minio/metrics/v3/tables
A scrape job using the root /minio/metrics/v3 endpoint satisfies the above requirement.
Replica catalog scanner stalled
This alert triggers when the replica catalog scanner has not completed a cycle recently. A stalled scanner means the replica catalog stops advancing, so the disaster-recovery site falls silently further behind the primary.
Because only the cluster leader runs the scanner, aggregate across the servers of one site rather than alerting per server.
The > 0 comparison excludes the nodes and sites that report a zero timestamp because they do not run the scanner, including the primary site.
These metrics carry only a server label, so they do not identify which deployment they came from.
If one Prometheus scrapes more than one site, a bare max() or sum() merges them: a healthy replica’s recent timestamp hides a stalled one, and a failure alert loses the site it came from.
Aggregate by whichever scrape label identifies the site in your setup — commonly job or cluster — and carry that label into the annotations, as below.
If each site has its own Prometheus, the bare aggregation is fine.
The threshold below suits the default replica_catalog_scanner_interval of 1m.
Raise it if you have set a longer interval.
alert: TablesReplicaCatalogScannerStalled
expr: time() - max by (job) (minio_tables_catalog_scanner_last_run_timestamp > 0) > 900
for: 5m
labels:
severity: critical
annotations:
summary: "AIStor Tables replica catalog scanner on {{ $labels.job }} has not completed a cycle in {{ $value | humanizeDuration }}"
impact: "The replica catalog is no longer advancing, so the DR site is falling behind the primary."
action: "Confirm the scanner is enabled with 'mc admin config get ALIAS tables'. Check the cluster leader's logs and run 'mc table replicate status' against the replica."
This alert has two limits.
It cannot detect a scanner that never started, because a site that has completed no cycle reports no timestamp above zero.
Confirm a new replica is scanning with mc table replicate status before you rely on this alert.
It also keeps firing on a site that was promoted by failover. Promotion turns the scanner off, but the last timestamp it reported stays above zero, so the expression keeps climbing on a site that is now correctly a primary. Silence this alert for the promoted site as part of your failover procedure.
Replica catalog scanner cycle failures
This alert triggers when catalog scanner cycles are failing.
Individual failures are retried on the next cycle, so this rule requires more than one failed cycle in the window: > 1 rather than > 0, which would fire on a single retryable failure.
The range must also be wide enough to hold two failed-cycle increments, or the rule can never reach its threshold.
Errors increment once per failed cycle, so the range has to span at least two cycles of replica_catalog_scanner_interval, and comfortably more to avoid sitting on the boundary.
Two hours is comfortable at the default 1m interval and at anything up to roughly 30m.
If you run a longer interval, widen the range to at least four times it — and update the last 2h wording in the summary annotation to match.
alert: TablesReplicaCatalogScannerErrors
expr: sum by (job) (increase(minio_tables_catalog_scanner_cycle_errors_total[2h])) > 1
for: 15m
labels:
severity: warning
annotations:
summary: "AIStor Tables catalog scanner on {{ $labels.job }}: {{ $value | humanize }} failed cycles in the last 2h"
impact: "The replica catalog may be advancing slowly or not at all."
action: "Check the cluster leader's logs for catalog scanner errors. Trace the scanner with 'mc admin trace'."
Both alerts require the Prometheus scraping configuration capture the metrics provided in the following API endpoint(s):
/minio/metrics/v3/tables
A scrape job using the root /minio/metrics/v3 endpoint satisfies the above requirement.