Kubernetes
Mount a MinIO AIStor Memory Bucket into a pod with the AIStor Memory CSI node driver — the workload pod gains no privilege, because kubelet mounts on the node.
Status — v1 driver, published image, no end-to-end CI
The driver image is published for each release and the driver is in v1. There is no automated end-to-end test that mounts through a live cluster yet, so validate the driver on a staging cluster before you depend on it in production. The driver has no Controller service: it mounts a Memory Bucket that already exists and never creates one.
A pod is disposable. Anything an agent writes to the container filesystem disappears when the pod does. The AIStor Memory CSI node driver gives that pod a durable workspace instead: it mounts a Memory Bucket as an ordinary directory, so everything written there persists to AIStor on your own storage with your own keys. The next pod mounts the same Memory Bucket and starts with what the last one wrote.
Why the CSI driver rather than mounting inside the pod
You can run aimem inside a pod, as the OpenShift
page shows. That pod needs CAP_SYS_ADMIN, privilege escalation for the setuid
mount helper, and access to /dev/fuse. A restrictive PodSecurity policy or an
OpenShift SCC blocks all three, and granting them to a pod that runs
model-authored code widens what that code can reach.
The CSI driver moves the mount off the pod. kubelet asks the driver to mount on the node before the container starts, and the container then sees a plain directory. The workload pod needs no capabilities at all:
securityContext:
allowPrivilegeEscalation: false
runAsNonRoot: true
runAsUser: 1000
capabilities:
drop: ["ALL"]
seccompProfile:
type: RuntimeDefaultThe privilege moves to one DaemonSet you install and audit once, instead of every workload pod.
What your nodes need
The mount runs on the node under systemd, so each node must provide:
- systemd, with PID 1 visible to the driver pod. The DaemonSet sets
hostPID: trueand usesnsenterto reach the host's systemd. Minimal hosts that do not run systemd, such as Talos, are out of scope in v1. /dev/fusepresent on the node.- A kubelet directory mounted with
Bidirectionalpropagation, so a mount made on the node reaches the workload pod. The DaemonSet below sets this.
You also need a Memory Bucket that already exists, and credentials that can mount it.
aimem bucket credentials <name> mints scoped credentials for one Memory Bucket.
The endpoint must resolve from the node
This is the mistake that costs the most time. Because the mount executes on the
node, endpointUrl has to resolve and connect from the node, not from the
pod.
A Kubernetes Service name does not work. The node sits outside the cluster
network and does not use cluster DNS, so http://aistor:9000 fails with a
transport error while a pod on the same cluster reaches it without trouble. Use
one of these instead:
- A Service ClusterIP, which kube-proxy programs on the node.
- Any address routable from the host. An external AIStor endpoint is the normal production case.
Install the driver
Pin the image to the release you are deploying. Do not use :latest: this
container performs the mount every workload on the node depends on, so a
floating tag lets a node restart change the filesystem implementation under
running pods.
Save this as aimem-csi.yaml, replace RELEASE_TAG with your release, and
apply it:
apiVersion: storage.k8s.io/v1
kind: CSIDriver
metadata:
name: csi.aimem.min.io
spec:
# Node-only driver: no Controller service, nothing to attach, so kubelet
# must not wait on an external-attacher.
attachRequired: false
podInfoOnMount: true
# kubelet calls the driver again while a volume is mounted, and each call
# brings the mount's credentials up to date.
requiresRepublish: true
volumeLifecycleModes:
- Persistent
- Ephemeral
fsGroupPolicy: None
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: aimem-csi-node-sa
namespace: kube-system
---
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: aimem-csi-node
namespace: kube-system
labels:
app: aimem-csi-node
spec:
selector:
matchLabels:
app: aimem-csi-node
template:
metadata:
labels:
app: aimem-csi-node
spec:
serviceAccountName: aimem-csi-node-sa
# Lets the driver nsenter into the host namespaces and launch aimem
# under the host's systemd.
hostPID: true
priorityClassName: system-node-critical
tolerations:
- operator: Exists
initContainers:
# Stage the aimem binary onto the host: the mount runs on the host,
# not in this pod.
- name: install-aimem
image: quay.io/minio/aistor/aimem-csi-driver:RELEASE_TAG
command:
- sh
- -c
- install -D -m 0755 /usr/local/bin/aimem /host/opt/aimem/bin/aimem
volumeMounts:
- name: aimem-host-bin
mountPath: /host/opt/aimem/bin
containers:
- name: aimem-csi-driver
image: quay.io/minio/aistor/aimem-csi-driver:RELEASE_TAG
args:
- --endpoint=$(CSI_ENDPOINT)
- --node-id=$(CSI_NODE_ID)
env:
- name: CSI_ENDPOINT
value: unix:///csi/csi.sock
- name: CSI_NODE_ID
valueFrom:
fieldRef:
fieldPath: spec.nodeName
- name: RUST_LOG
value: info
securityContext:
privileged: true
livenessProbe:
httpGet:
path: /healthz
port: healthz
initialDelaySeconds: 10
periodSeconds: 60
timeoutSeconds: 3
ports:
- name: healthz
containerPort: 9808
protocol: TCP
volumeMounts:
- name: plugin-dir
mountPath: /csi
- name: kubelet-dir
mountPath: /var/lib/kubelet
# Bidirectional so the host-side mount propagates into workload
# pods.
mountPropagation: Bidirectional
- name: aimem-host-bin
mountPath: /opt/aimem/bin
- name: node-driver-registrar
image: registry.k8s.io/sig-storage/csi-node-driver-registrar:v2.13.0
args:
- --csi-address=/csi/csi.sock
- --kubelet-registration-path=/var/lib/kubelet/plugins/csi.aimem.min.io/csi.sock
volumeMounts:
- name: plugin-dir
mountPath: /csi
- name: registration-dir
mountPath: /registration
- name: liveness-probe
image: registry.k8s.io/sig-storage/livenessprobe:v2.16.0
args:
- --csi-address=/csi/csi.sock
- --health-port=9808
volumeMounts:
- name: plugin-dir
mountPath: /csi
volumes:
- name: plugin-dir
hostPath:
path: /var/lib/kubelet/plugins/csi.aimem.min.io/
type: DirectoryOrCreate
- name: registration-dir
hostPath:
path: /var/lib/kubelet/plugins_registry/
type: Directory
- name: kubelet-dir
hostPath:
path: /var/lib/kubelet
type: Directory
- name: aimem-host-bin
hostPath:
path: /opt/aimem/bin
type: DirectoryOrCreatekubectl apply -f aimem-csi.yaml
kubectl -n kube-system rollout status daemonset/aimem-csi-nodeA node-only driver that provisions nothing needs no cluster-scoped permissions, which is why the ServiceAccount above carries no Role. It exists to give the DaemonSet a stable identity.
Choose a volume style
Both styles mount a Memory Bucket that already exists. They differ in who authors the mount and how long it lives.
| Style | Who authors it | Lives as long as | Use it when |
|---|---|---|---|
| Inline ephemeral | Whoever creates the pod | The pod | One mount per unit of work — agent sandboxes, CI jobs |
| PersistentVolume | A cluster admin | The PV | A long-lived, shared workspace, or when only admins may name buckets |
Inline ephemeral leaves nothing behind for a controller to garbage-collect. A
PersistentVolume puts the mount's settings under admin control, which matters
because localDir is accepted only there.
Mount with an inline ephemeral volume
The pod names the Memory Bucket directly. bucketName is required here: kubelet
generates the volume handle from the pod, so the fallback that a
PersistentVolume uses would name a bucket that does not exist, and the driver
rejects the request rather than failing later.
apiVersion: v1
kind: Secret
metadata:
name: aimem-inline-credentials
type: Opaque
stringData:
# Scope these to one Memory Bucket where you can:
# `aimem bucket credentials <name>`.
accessKeyID: REPLACE_ME
secretAccessKey: REPLACE_ME
# sessionToken: REPLACE_ME # required for STS credentials
---
apiVersion: v1
kind: Pod
metadata:
name: aimem-inline-workspace
spec:
restartPolicy: Never
containers:
- name: workload
image: ubuntu:24.04
command: ["bash", "-lc"]
args:
- |
set -euo pipefail
ls -la /workspace
# This write is durable: it lands in the Memory Bucket, not in the pod.
printf 'hello from %s\n' "$(hostname)" > /workspace/probe.txt
sync
volumeMounts:
- name: workspace
mountPath: /workspace
securityContext:
allowPrivilegeEscalation: false
runAsNonRoot: true
runAsUser: 1000
capabilities:
drop: ["ALL"]
seccompProfile:
type: RuntimeDefault
volumes:
- name: workspace
csi:
driver: csi.aimem.min.io
volumeAttributes:
bucketName: my-project
# Mount a sub-prefix instead of the whole Memory Bucket. Leading and
# trailing slashes are optional.
# prefix: workspaces/alice/app-7/
endpointUrl: https://aistor.example.com:9000
nodePublishSecretRef:
name: aimem-inline-credentialsMount with a PersistentVolume
An admin defines the volume, and the pod claims it. volumeHandle names the
Memory Bucket when bucketName is absent.
apiVersion: v1
kind: PersistentVolume
metadata:
name: aimem-my-project
spec:
capacity:
storage: 1Ti # Ignored by the driver; Kubernetes requires a value.
accessModes: ["ReadWriteMany"]
persistentVolumeReclaimPolicy: Retain
storageClassName: "" # Static: bind by claim, not by class.
csi:
driver: csi.aimem.min.io
volumeHandle: my-project
volumeAttributes:
endpointUrl: https://aistor.example.com:9000
# Staging and read cache on the node. PersistentVolume only.
localDir: /var/lib/aimem/staging/my-project
nodePublishSecretRef:
name: aimem-credentials
namespace: kube-systemVolume attributes
Set these under volumeAttributes.
| Attribute | Effect |
|---|---|
bucketName | The Memory Bucket to mount. Defaults to volumeHandle; required on an inline volume. |
prefix | Mount a sub-prefix, as bucket/prefix/. Surrounding slashes are normalised. |
endpointUrl | The AIStor endpoint. Must resolve from the node. |
region | The region to sign requests for. |
agent | The agent identity AIStor stamps on writes. |
metadataTtl | How long metadata stays cached. |
tlsCaFile | A CA bundle path on the node, not in the pod. |
localDir | Staging and read cache on the node. PersistentVolume only. |
readOnly | Mount read-only. Also honoured from the CSI request flag. |
Attributes the driver rejects
uid, gid, allowOther, allowRoot, and virtualHostStyle are rejected
with InvalidArgument naming the attribute. The aimem CLI has no matching
flag, so there is nothing to translate them to. The driver refuses the mount
rather than ignoring the request, because a volume that asks for allowOther
and comes up without it looks like it worked while behaving differently from
what it declares.
localDir is rejected on an inline ephemeral volume, and only there. It names a
host path, and an inline volume's attributes are authored by whoever creates the
pod — honouring it would let a pod stage into another volume's directory, or
anywhere else on the node. A PersistentVolume's attributes are an
administrator's, so it is accepted there. It must still be absolute and free of
...
Credentials
The driver writes each mount's credentials to a file that only root can read,
and gives aimem the file's path in AIMEM_CREDENTIALS_FILE. The credentials
never appear in the mount's environment or command line, so systemctl show
does not reveal them. The driver deletes the file when the volume is unmounted.
Because the CSIDriver sets requiresRepublish, kubelet calls the driver again
while a volume is mounted. Each call brings the file up to date, and aimem
switches to the new credentials within a few seconds. A mount can therefore
outlive the session it started with, provided each new session reaches it
before the old one expires.
kubelet calls periodically, not at a fixed deadline, so give sessions enough lifetime to cover the time between calls and any delay. A session that expires before the next call leaves the mount unable to reach the bucket until that call arrives. Sessions of an hour or more leave ample margin.
Each volume gets its credentials from one of two sources. When a volume has a Secret, the Secret is used.
From a Secret
Name the Secret with nodePublishSecretRef. It holds these keys:
| Secret key | Required |
|---|---|
accessKeyID | Yes |
secretAccessKey | Yes |
sessionToken | Only for temporary sessions |
To rotate credentials, update the Secret. A running mount picks up the new keys on kubelet's next call, with no remount.
From the pod's service account
A volume with no Secret can get a session by exchanging the pod's service-account token. Set it up in four steps.
-
Run an endpoint that accepts this request and returns a session:
POST <url> Authorization: Bearer <service-account token> {"bucket": "...", "prefix": "..."} 200 {"access_key_id": "...", "secret_access_key": "...", "session_token": "...", "expiration": "<RFC 3339>"} -
Add these arguments to the
aimem-csi-drivercontainer:args: - --token-exchange-url=https://sessions.example.com/exchange - --token-audience=aimemThe environment variables
AIMEM_CSI_TOKEN_EXCHANGE_URLandAIMEM_CSI_TOKEN_AUDIENCEwork too. -
Ask kubelet for a token of that audience in the
CSIDriverspec:tokenRequests: - audience: aimem expirationSeconds: 3600 -
Leave
nodePublishSecretRefoff the volume.
The driver exchanges the token again once half the session's lifetime has
passed. A volume with neither a Secret nor a configured exchange fails to
publish, with an error naming the missing nodePublishSecretRef.
Where the files live
The driver keeps credential files in --credentials-dir
(AIMEM_CSI_CREDENTIALS_DIR), which defaults to
/var/lib/kubelet/plugins/csi.aimem.min.io/credentials. The mount runs on the
host and reads the file there, so the directory must have the same path on the
host and in the driver pod. The manifest above mounts /var/lib/kubelet at its
own path, which meets this.
How a mount comes up and goes away
On NodePublishVolume the driver launches aimem on the host under a transient
systemd unit, then polls /proc/self/mountinfo until the mount appears. Because
host PID 1 owns the mount, restarting or upgrading the driver DaemonSet does not
tear down live mounts.
When kubelet calls NodePublishVolume again for a live mount, the driver only
updates the mount's credentials file. It does not relaunch the mount.
On NodeUnpublishVolume it stops the unit, unmounts, removes the target
directory, and deletes the credentials file.
Limitations in v1
- No dynamic provisioning. There is no Controller service, so the driver never creates a Memory Bucket. Both volume styles name one that already exists.
- systemd nodes only. See What your nodes need.