Skip to main content
Goal: query live connectors into a remote Python kernel and analyze there — the data never lands on your machine.

1.1 · See your connectors, then look before you query

Prompt
You’ll see: the org’s connectors, then real table names and sample rows. Raw SQL executes verbatim in the connector’s own dialect — check refinery connector db get [id] and write that dialect; there is no translation layer.

1.2 · Query into a sandbox you named

Prompt
You’ll see: the result lands as a pandas dataframe inside the named sandbox; the next exec python sees it directly. Kernel state persists across calls; files under /sandbox/files survive restarts, variables don’t.
⚠️ Name your sandbox after the task — never share “default” — Two sessions that both use the default name share one kernel: variables silently overwrite each other, and one sandbox rm destroys the other’s state. --sandbox churn-analysis costs nothing and saves an afternoon.

1.3 · The two disciplines

  • Aggregate in place. Never pull large result sets to your terminal — compute in the sandbox and print only the summary. Memory is bounded; avoid df.copy().
  • Python never fetches data. Retrieval order is fixed: search the ontology for an existing definition → run a governed .tql → only then raw SQL. Python is for post-processing and charting.

✅ Checkpoint

  • A dataframe loaded via --as is visible from a follow-up exec python in your named sandbox
  • You checked the connector’s dialect before writing SQL for it
  • You can recite the retrieval order (ontology → .tql → raw SQL) and where Python fits