> ## Documentation Index
> Fetch the complete documentation index at: https://docs.textql.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Troubleshooting

> Get logs and diagnose problems in a self-hosted TextQL deployment

<Note>
  This guide assumes `kubectl` is already configured and pointed at the right cluster and namespace.
</Note>

## How to share information with the TextQL team

### TextQL Doctor

Use the TextQL Doctor utility (shell script) to quickly generate an encrypted file to share with TextQL. The file contains logs, cluster configuration, and resource status.

```bash theme={null}
# requires `kubectl`
./textql-doctor.sh -n <namespace>          # use the k8s namespace here
./textql-doctor.sh -n <namespace> --auto   # to auto approve all commands
```

`textql-doctor.sh` ships with your Helm chart (versions 1.2.3 and later, under `textql-charts/`). If you don't have access to the chart, or run an older version, you can request the script from the TextQL team.

### Alternative

If you can't run TextQL Doctor, provide the output of these commands:

```bash theme={null}
# 1. Pod snapshot
kubectl get pods

# 2. compute-engine logs (the main one)
kubectl logs deploy/compute-engine > tql-compute-engine.log

# 3. If requested, fetch logs from other components as well
kubectl logs deploy/ontology > tql-ontology.log

# 4. Describe output for any pod that is not Ready
kubectl describe pod <pod-name> > tql-describe-<pod>.txt

# 5. (Optional) Output of Helm status, in the event of a failed upgrade/install
helm status <release-name>
```

Also include:

* The exact error message from the browser **Console** or **Network** tabs if the issue is reproducible in the UI.
* Your `values.yaml` file and/or the relevant attributes in that file. Make sure you redact any private information like tokens.

In the following sections, you'll find a more detailed guide on how to troubleshoot the deployment yourself, so you can identify issues and relevant logs before reaching out.

## 1. What runs in the cluster

TextQL deploys a small set of Deployments into a single namespace. Knowing which one to look at is the biggest single win when triaging.

| Workload | What it does | When to suspect it |
| - | - | - |
| `web` | User frontend, auth, RPC entrypoint. | Login spinning, blank pages, 400s on `/rpc/*`, UI issues. |
| `compute-engine` | Orchestrates chats, sandboxes, connectors, LLM calls. **This is where most things you'll debug live.** | Chats, playbooks, dashboards, tool calls or connector failures, sandbox issues, identity or setting misconfigurations, API errors, database-related bugs. |
| `sandbox-proxy` | TLS-terminating proxy in front of Python sandboxes (mTLS via CA cert). | Sandbox runs fail, "Failed to parse CA certificate". |
| `ontology` | Natural-language to SQL, ontology execution. | Metrics query failures, `FatalError` on ontology queries. |
| `valkey` | Redis-compatible cache. | Slow auth, stale results, intermittent 5xx. |
| `sandbox pods` | Per-execution Python pods, ephemeral. | Python/SQL cell failures, OOMKilled, sandbox timeouts, `No space left on device` (see [Sandbox Storage](/helm-charts/sandbox-storage)). |
| `oathkeeper` | API gateway that sits in front of all traffic. Handles route matching and authentication. | SCIM requests failing. SSO/OIDC login loops. A route returning 404 or categories of endpoints unreachable. Auth working for some methods but not others. |

## 2. Inspect the environment

```bash theme={null}
# Default to the TextQL namespace so you don't need -n every time
kubectl config set-context --current --namespace=<namespace>

# Pods in the namespace
kubectl get pods

# Inspect a specific deployment (replicas, image, conditions, events)
kubectl describe deployment <deployment-name>

# This can surface variables that are not correctly set, you can filter those
kubectl describe deployment <deployment-name> | grep "SANDBOX_AUTH_KEY"

# Per-pod CPU and memory
kubectl top pods
```

## 3. Getting logs from compute-engine

This is the section you'll come back to. Most issues you'll debug live in `compute-engine`.

### 3.1 Print the logs

```bash theme={null}
kubectl logs deploy/compute-engine --all-containers -f
```

This command tails the live log stream for all containers in the `compute-engine` pod. Add `--previous` if the pod just restarted and you want what it logged before crashing.

### 3.2 Grep for specific errors

Pipe the log output through `grep` to find what you need.

```bash theme={null}
kubectl logs deploy/compute-engine --all-containers --since=30m | grep -iE 'error|fatal|panic'
```

The TextQL team can point you at specific filters that you can apply in this step depending on your issue.

### 3.3 When the pod isn't healthy

If pods are not Ready or restarting, `describe` tells you why.

```bash theme={null}
kubectl describe pod -l app=compute-engine

# Force a rollout (picks up updated secrets / configmaps)
kubectl rollout restart deploy/compute-engine
kubectl rollout status deploy/compute-engine
```

## 4. Debugging network problems from inside the cluster

When you suspect a connectivity issue (to the DB, object storage, an LLM provider, or another in-cluster service), you can attach a debug container to the pod with networking tools. The `nicolaka/netshoot` image bundles `curl`, `dig`, `nslookup`, `nc`, `tcpdump`, `mtr`, and others.

```bash theme={null}
# Attach a netshoot container to the compute-engine pod, sharing its network namespace
kubectl debug -it <pod-name> --image=nicolaka/netshoot

# Once inside, you're on the same network as compute-engine. Try:

# DNS resolution
nslookup <db-endpoint>
dig +short <db-endpoint>

# TCP reachability to the DB
nc -vz <db-endpoint> 5432

# HTTPS to an external provider
curl -v https://api.anthropic.com/v1/messages -o /dev/null

# Reach another in-cluster service
curl -v http://ontology.<namespace>.svc.cluster.local:5000/health

# See the routing table
ip route

# Watch live traffic on a port
tcpdump -i any -nn port 5432
```

You can also spin up a standalone netshoot pod (not attached to anything) when you want a generic in-cluster network shell:

```bash theme={null}
kubectl run netshoot --rm -it --image=nicolaka/netshoot -- /bin/bash
```

Common things to confirm:

* **DNS works:** if `nslookup` fails, cluster DNS is misconfigured.
* **TCP port is open:** `nc -vz` succeeding rules out security-group / network-policy issues.
* **TLS handshake completes:** `curl -v` can show you related information.
* **In-cluster service DNS resolves:** `<service>.<namespace>.svc.cluster.local`.

## 5. Logs from the browser (DevTools)

Many issues show up in the user's browser, which provides an easier view. Right-click anywhere on the page → **Inspect** (or `Cmd+Option+I` on Mac, `Ctrl+Shift+I` on Windows/Linux).

### 5.1 Console tab: JavaScript errors and stack traces

The Console tab shows everything the frontend logged: unhandled promise rejections, RPC errors with full server-side detail, and toast errors that disappeared too fast to read.

**What to look for:**

* **Red lines are errors.** The first red line is usually the actionable one.
* **HTTP error codes** identify the type of error, e.g. 404 for a resource not found, or 500 for an internal server error.
* Examples:
  * `ConnectError: [code] message`: a Connect-RPC failure from `compute-engine`. The bracketed code (`[unavailable]`, `[invalid_argument]`, `[unauthenticated]`, `[internal]`) tells you the category.
  * `Failed to fetch` / CORS errors: the load balancer is misconfigured, or the object storage bucket CORS policy doesn't include your domain in `AllowedOrigins`.
  * `WebSocket connection failed`: the load balancer is dropping long-lived connections (idle timeout too low), or `/rpc/public` isn't routed to `compute-engine`.

**Actions to improve the output:**

1. Toggle **Preserve log** so you don't lose errors on page reload.
2. Turn on timestamps (Settings cog → Console).
3. Right-click any line → **Save as** to capture the full Console output.
4. Filter by level (Errors only) when there's a lot of noise.

### 5.2 Network tab: what the app actually asked the backend

When chats hang, files won't upload, or an action "does nothing," the Network tab tells you whether the request left the browser, where it went, and what came back.

**Setup once per session:**

1. Open DevTools → **Network** tab.
2. Toggle **Preserve log**.
3. Toggle **Disable cache** while DevTools is open.
4. Filter by **Fetch/XHR** to hide static assets, shows only API calls.
5. Reproduce the issue.

**Useful Network filter strings:**

```
/rpc/               # all Connect-RPC calls to compute-engine
/v1/                 # versioned REST endpoints
/sandbox/            # sandbox-proxy traffic (file uploads, exec)
/asset/              # presigned URLs for assets
status-code:500      # any 5xx
larger-than:1M       # oversized uploads/downloads
```

For any one request, the most useful sub-tabs are:

* **Payload**: what the frontend actually sent (chat ID, org ID, etc.).
* **Response**: the server's response body. Connect-RPC errors include structured JSON with `code` and `message`. Copy verbatim when reaching out for support.
* **Timing**: where time was spent, and possible timeouts.

## 6. Helm commands

If you installed TextQL via Helm and the installation or upgrade failed (see [Upgrading](/helm-charts/upgrading) for the upgrade checklist and rollback), review the output of the installation process. This can surface errors like:

* Secrets not being created
* Timeouts on pending upgrades
* Attributes missing in `values.yaml`

```bash theme={null}
# To get the release name
helm list -A

# Get the particular status of your installation
helm status <release-name>

# Errors related to previous installations
helm history <release-name>
```

These commands are generally useful even if you use Flux or ArgoCD on top of Helm.

## 7. SCIM: Reference for Self-Hosted / VPC Deployments

This section extends [SCIM Provisioning](/core/admin/scim) with the endpoints reference, operational commands, and Helm values relevant to operators of a self-hosted deployment with `kubectl` access to the cluster.

### Endpoints Reference

All SCIM endpoints are served under `/scim/v2/`:

| Method | Path | Description |
| - | - | - |
| `GET` | `/scim/v2/Users` | List users |
| `POST` | `/scim/v2/Users` | Provision a user |
| `GET` / `PUT` / `PATCH` / `DELETE` | `/scim/v2/Users/:id` | Manage a user |
| `GET` | `/scim/v2/Groups` | List groups |
| `POST` | `/scim/v2/Groups` | Create a group |
| `GET` / `PUT` / `PATCH` / `DELETE` | `/scim/v2/Groups/:id` | Manage a group |
| `GET` | `/scim/v2/ServiceProviderConfig` | SCIM capability discovery |
| `GET` | `/scim/v2/ResourceTypes` | Resource type definitions |
| `GET` | `/scim/v2/Schemas` | Schema definitions |

### Operational Commands

Replace `<namespace>` with your deployment namespace (typically `textql`).

#### Get compute-engine logs

SCIM operations are handled by the `compute-engine` pod. All SCIM auth checks, provisioning events, and errors are logged there.

```bash theme={null}
# Stream live logs
kubectl logs -f deployment/compute-engine -n <namespace>

# Search for all SCIM-related log lines
kubectl logs deployment/compute-engine -n <namespace> | grep -i scim

# Search for SCIM token validation events
kubectl logs deployment/compute-engine -n <namespace> | grep "SCIM token check"

# Search for SCIM auth failures
kubectl logs deployment/compute-engine -n <namespace> | grep "SCIM auth failed"

# Search for group-level operations
kubectl logs deployment/compute-engine -n <namespace> | grep -i "scim.*group\|group.*scim"

# Search for provisioning errors
kubectl logs deployment/compute-engine -n <namespace> | grep -E "SCIM.*(error|failed|ERROR)"

# Tail the last N lines and filter
kubectl logs deployment/compute-engine -n <namespace> --tail=10000 | grep -i scim
```

If you see `SCIM token check: valid` but no subsequent group operation logs, the group sync requests are being dropped before reaching the application (likely a network issue).

#### Get the Oathkeeper configmap

The Oathkeeper configmap controls how SCIM requests are routed and authenticated. If SCIM is failing entirely for a VPC deployment, this is the first thing to inspect.

```bash theme={null}
kubectl get configmap oathkeeper-config -n <namespace> -o yaml
```

What to look for in the configmap:

1. A rule with `"id": "scim:v2:authenticated"` must exist, matching `<http|https>://<[^/]+>/scim/v2/<.*>`.
2. The frontend catch-all rule (`web:frontend:authenticated`) must have `scim/` in its negative lookahead exclusion pattern, e.g. `(?!health|v1/|rpc/|...|scim/|...)`. If `scim/` is absent, Oathkeeper will intercept SCIM requests and apply browser-based OIDC auth instead.

#### Get the Oathkeeper deployment and check its version

```bash theme={null}
# Check Oathkeeper pod status
kubectl get pods -n <namespace> -l app=oathkeeper

# Check which image version Oathkeeper is running
kubectl get deployment oathkeeper -n <namespace> -o jsonpath='{.spec.template.spec.containers[0].image}'

# View Oathkeeper logs (useful if the SCIM rule itself is malformed)
kubectl logs deployment/oathkeeper -n <namespace> | grep -i "scim\|error\|rule"
```

#### Check and adjust the SCIM rate limit

```bash theme={null}
# Check the current rate limit configured for the compute engine
kubectl get deployment compute-engine -n <namespace>  # -> review the "SCIM_RATE_LIMIT" env variable

# Edit the deployment to set a higher limit (e.g. 1000 req/min)
kubectl set env deployment/compute-engine -n <namespace> SCIM_RATE_LIMIT=1000
kubectl rollout status deployment/compute-engine -n <namespace>
```

The default is 500 requests per minute. If this variable is absent, the default applies.

#### Restart compute-engine (after a configmap or env var change)

```bash theme={null}
kubectl rollout restart deployment/compute-engine -n <namespace>
kubectl rollout status deployment/compute-engine -n <namespace>
```

### Helm Values Reference

These `values.yaml` keys control SCIM and OIDC behavior for self-hosted deployments. They're set at deployment time and take effect after a `helm upgrade`.

| Key | Type | Default | Description |
| - | - | - | - |
| `compute.scimRateLimit` | integer | `500` | Per-org SCIM request rate limit (requests per minute). Exposed as the `SCIM_RATE_LIMIT` env var on the compute engine pod. Raise this for large initial provisioning runs. |
| `oathkeeper.enabled` | boolean | `false` | Enable Oathkeeper for load balancer-level authentication. Must be `true` for VPC deployments using OIDC or SCIM bearer token auth at the edge. |
| `oathkeeper.enforceAuth` | boolean | `false` | Enforce authentication at the Oathkeeper level: routes the frontend through Oathkeeper, strips anonymous handlers from rules, and redirects unauthenticated browser requests to `/oidc/authorize`. When `false`, Oathkeeper still routes and validates SCIM tokens but does not enforce browser-level auth. |
| `global.auth.singleOIDCTenant` | boolean | `false` | Use a single env-based OIDC provider for all authentication. Requires `oathkeeper.enabled=true`, OIDC env vars on compute and web, and `web.publicApi` set. |
| `global.auth.singleOidcIdleLogoutAfter` | string | `""` | Revoke browser auth after this idle period with no requests through Oathkeeper. Format: `<digits>s\|m\|h` (e.g. `45m`, `12h`). Empty means disabled. Requires `singleOIDCTenant`. |
| `global.auth.oidc.providerType` | string | `""` | OIDC provider type. Accepted values: `Microsoft OIDC`, `Ping OIDC`, `Okta OIDC`. |
| `global.auth.oidc.issuerUrl` | string | `""` | OIDC issuer URL. Example: `https://login.microsoftonline.com/{tenant}/v2.0` for Entra ID, `https://{your-domain}.okta.com` for Okta. |
| `global.auth.oidc.clientId` | string | `""` | OIDC application client ID. |
| `global.auth.oidc.clientSecret` | string | `""` | **(Secret)** OIDC application client secret. |
| `global.auth.oidc.scopes` | string | `""` | Space-separated OIDC scopes. Defaults to `openid profile email` if not set. |
| `global.auth.oidc.attributeMapping` | string | `""` | Optional JSON mapping of OIDC claims to TextQL user attributes. Used when the IdP uses non-standard claim names. |
| `global.auth.oidc.emailVerification` | boolean | `true` | Set to `false` to allow unverified OIDC emails. Only disable if your IdP does not mark emails as verified. |

Corresponding environment variables (set automatically from values by the Helm chart):

| Env var | Source value |
| - | - |
| `SINGLE_OIDC_TENANT` | `global.auth.singleOIDCTenant` |
| `SINGLE_OIDC_IDLE_LOGOUT_AFTER` | `global.auth.singleOidcIdleLogoutAfter` |
| `OIDC_PROVIDER_TYPE` | `global.auth.oidc.providerType` |
| `OIDC_ISSUER_URL` | `global.auth.oidc.issuerUrl` |
| `OIDC_CLIENT_ID` | `global.auth.oidc.clientId` |
| `OIDC_CLIENT_SECRET` | `global.auth.oidc.clientSecret` (from `oidc-secrets` k8s secret) |
| `OIDC_SCOPES` | `global.auth.oidc.scopes` |
| `OIDC_ATTRIBUTE_MAPPING` | `global.auth.oidc.attributeMapping` |
| `OIDC_EMAIL_VERIFICATION` | `global.auth.oidc.emailVerification` |
| `SCIM_RATE_LIMIT` | `compute.scimRateLimit` |

## Getting Support

If you're still stuck after working through this guide, contact [support@textql.com](mailto:support@textql.com) with the output of TextQL Doctor (or the commands under [Alternative](#alternative)) attached.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.