Skip to main content
Every Python sandbox in a self-hosted deployment writes to one shared volume, sandbox-files-pvc, which the chart mounts into compute-engine at /mnt/sandbox-data. This guide covers how much space that volume needs, how retention keeps it bounded, and how to recover when it fills up. Chart values below are set in the values file you pass to helm upgrade (see Installation).

What lives on the volume

The volume is shared by every sandbox in the deployment, so when it fills up, every thread fails to write, not only the one that pushed it over. Threads report No space left on device, and dashboards fail with NameError: name 'dashboard_<source>' is not defined because their data sources could not be saved.

Capacity planning

Ontology size has little effect on storage. The library directory is one small shared copy. What drives growth is the number of active sandboxes and how much data each one pulls into Python. Every dataframe in a thread’s Python session is saved into its workspace snapshot, so a thread that only queries aggregates uses a few MB, while one that loads millions of rows can use several GB. Usage levels off instead of growing forever, because retention deletes idle sandbox directories (see Retention). At steady state the volume holds roughly:
To measure your own numbers, run this from the compute-engine pod. On a large volume du can take several minutes.
Size the volume for two to three times the current steady state so growth in users and scheduled playbooks does not force another resize soon. Prefer growing the volume over shortening retention: a larger volume costs disk, while shorter retention makes more old threads restore from backup and drops the files they wrote. Whatever size you pick, alert on free space so you hear about it before users do. On GKE, alert on the Filestore metric file.googleapis.com/nfs/server/free_bytes_percent, for example below 20%. On AKS, alert on the file share’s used capacity against its quota.

Retention

Two settings control how long sandbox data is kept. compute.worker.directoryRetentionDays (Helm, default 15, or 30 in chart 1.3.25 and earlier) controls the shared volume. A daily job deletes any sandbox-worker-* directory whose last activity is older than this. Activity is measured from the workspace snapshot’s last write (or the directory’s, when there is no snapshot), not from when the sandbox was created, so a long-running thread that is still in use is kept. Workspace Retention (Settings → Security, default 30 days) controls a backup copy of each thread’s Python workspace in object storage, up to 5 GiB per thread. See Settings. When someone reopens a thread whose directory was deleted, the sandbox restores the Python workspace from the backup copy (if it is still within Workspace Retention) and re-attaches chat attachments, so the thread picks up where it left off after a slightly slower start. Files Ana wrote to disk that were not kept as artifacts, and workspaces larger than 5 GiB, are not restored. Dashboards reload their data sources the next time they start.
Lower directoryRetentionDays only when you cannot grow the volume. Keep Workspace Retention at or above it so a thread’s backup outlives its directory.

Platform differences

Expanding on GKE

Filestore Basic HDD has a 1 TiB minimum, and the Filestore driver rounds smaller requests up to it. With the default filesPvcSize of 128Gi you already have 1 TiB, so a new size only helps if it is larger than the volume’s current capacity. Check it first:
The chart creates the Filestore StorageClass with expansion turned off. To grow the volume:
  1. Allow expansion on the StorageClass. Its name is filestore-<namespace> unless you set sandbox.storageClassName.
  2. Set the new size and the expansion flag in your values file, then run helm upgrade.
  3. Confirm the new capacity with kubectl get pvc sandbox-files-pvc -n <namespace>. Filestore expands in place without downtime. It cannot shrink, so increase in steps you are confident you will use.

Keep the volume on uninstall

Dynamically provisioned volumes default to the Delete reclaim policy, so deleting the PVC (for example by uninstalling the release) can delete the underlying storage along with every sandbox and the context library. Set the policy to Retain on every platform:

Recovering from a full volume

Work through these in order. Freeing space ends the outage. Expanding the volume and checking retention keep it from coming back.
  1. Back up the volume. On GKE, find the Filestore instance backing the volume (kubectl get pvc sandbox-files-pvc -n <namespace> shows its name under VOLUME) in the GCP console (Filestore → Instances) and create an on-demand backup under Backups. On EKS, use AWS Backup for the EFS file system. On AKS, take a snapshot of the file share.
  2. Optionally keep a local copy of the context library.
  3. Record usage before deleting anything. Run the measurement commands from Capacity planning and save the output. These numbers are your best sizing input, because they show the volume at its peak.
  4. List idle sandbox directories. This uses the same activity rule as the daily cleanup job. Replace 14 with the number of idle days you are comfortable deleting.
  5. Delete only the listed directories.
    Never delete library, packages, or anything that does not start with sandbox-worker-. Deleting a sandbox that is in use breaks that thread or dashboard until it restarts.
  6. Reload affected work. Dashboards recover once space is free and the page is reloaded. Files that failed to write while the volume was full are gone, so rerun those steps in their threads.
  7. Prevent a repeat. Set the reclaim policy to Retain, grow the volume using the numbers from step 3 (see Platform differences), add a free-space alert, and review Retention.
Contact TextQL with the output from step 3 if you want help choosing a size.