sandbox-files-pvc, which the chart mounts into compute-engine at /mnt/sandbox-data. This guide covers how much space that volume needs, how retention keeps it bounded, and how to recover when it fills up. Chart values below are set in the values file you pass to helm upgrade (see Installation).
What lives on the volume
The volume is shared by every sandbox in the deployment, so when it fills up, every thread fails to write, not only the one that pushed it over. Threads report
No space left on device, and dashboards fail with NameError: name 'dashboard_<source>' is not defined because their data sources could not be saved.
Capacity planning
Ontology size has little effect on storage. Thelibrary directory is one small shared copy. What drives growth is the number of active sandboxes and how much data each one pulls into Python. Every dataframe in a thread’s Python session is saved into its workspace snapshot, so a thread that only queries aggregates uses a few MB, while one that loads millions of rows can use several GB.
Usage levels off instead of growing forever, because retention deletes idle sandbox directories (see Retention). At steady state the volume holds roughly:
compute-engine pod. On a large volume du can take several minutes.
file.googleapis.com/nfs/server/free_bytes_percent, for example below 20%. On AKS, alert on the file share’s used capacity against its quota.
Retention
Two settings control how long sandbox data is kept.compute.worker.directoryRetentionDays (Helm, default 15, or 30 in chart 1.3.25 and earlier) controls the shared volume. A daily job deletes any sandbox-worker-* directory whose last activity is older than this. Activity is measured from the workspace snapshot’s last write (or the directory’s, when there is no snapshot), not from when the sandbox was created, so a long-running thread that is still in use is kept.
Workspace Retention (Settings → Security, default 30 days) controls a backup copy of each thread’s Python workspace in object storage, up to 5 GiB per thread. See Settings.
When someone reopens a thread whose directory was deleted, the sandbox restores the Python workspace from the backup copy (if it is still within Workspace Retention) and re-attaches chat attachments, so the thread picks up where it left off after a slightly slower start. Files Ana wrote to disk that were not kept as artifacts, and workspaces larger than 5 GiB, are not restored. Dashboards reload their data sources the next time they start.
directoryRetentionDays only when you cannot grow the volume. Keep Workspace Retention at or above it so a thread’s backup outlives its directory.
Platform differences
Expanding on GKE
Filestore Basic HDD has a 1 TiB minimum, and the Filestore driver rounds smaller requests up to it. With the defaultfilesPvcSize of 128Gi you already have 1 TiB, so a new size only helps if it is larger than the volume’s current capacity. Check it first:
-
Allow expansion on the StorageClass. Its name is
filestore-<namespace>unless you setsandbox.storageClassName. -
Set the new size and the expansion flag in your values file, then run
helm upgrade. -
Confirm the new capacity with
kubectl get pvc sandbox-files-pvc -n <namespace>. Filestore expands in place without downtime. It cannot shrink, so increase in steps you are confident you will use.
Keep the volume on uninstall
Dynamically provisioned volumes default to theDelete reclaim policy, so deleting the PVC (for example by uninstalling the release) can delete the underlying storage along with every sandbox and the context library. Set the policy to Retain on every platform:
Recovering from a full volume
Work through these in order. Freeing space ends the outage. Expanding the volume and checking retention keep it from coming back.-
Back up the volume. On GKE, find the Filestore instance backing the volume (
kubectl get pvc sandbox-files-pvc -n <namespace>shows its name underVOLUME) in the GCP console (Filestore → Instances) and create an on-demand backup under Backups. On EKS, use AWS Backup for the EFS file system. On AKS, take a snapshot of the file share. -
Optionally keep a local copy of the context library.
- Record usage before deleting anything. Run the measurement commands from Capacity planning and save the output. These numbers are your best sizing input, because they show the volume at its peak.
-
List idle sandbox directories. This uses the same activity rule as the daily cleanup job. Replace
14with the number of idle days you are comfortable deleting. -
Delete only the listed directories.
Never delete
library,packages, or anything that does not start withsandbox-worker-. Deleting a sandbox that is in use breaks that thread or dashboard until it restarts. - Reload affected work. Dashboards recover once space is free and the page is reloaded. Files that failed to write while the volume was full are gone, so rerun those steps in their threads.
-
Prevent a repeat. Set the reclaim policy to
Retain, grow the volume using the numbers from step 3 (see Platform differences), add a free-space alert, and review Retention.