Skip to content

Datasets from Python

Versioning a graph needs no network. Putting rows somewhere a trial can read them does, so that half is a small HTTP client and not a Store.

import chatty_lab as cl
lab = cl.Platform.from_env() # CHATTY_API, CHATTY_TOKEN, CHATTY_WORKSPACE

Or explicitly:

lab = cl.Platform("https://chatty-lab.com", token, workspace=TEAM_ID)

The address may be the API root or the address of the site. workspace is the team whose things to speak about — omit it for your personal workspace. repr(lab) deliberately omits the token, so a notebook that prints the object does not print the secret.

The client has no dependencies: urllib and a bearer. A wheel installed to version a graph should not drag an HTTP library behind it.

ds = lab.new_dataset(
"medqa",
columns=[("input", "text"), ("reference", "text")],
description="USMLE questions, 500 first",
)

Four column types, and the widget follows from the type — the platform checks that they match, so this is not somewhere to improvise:

typeWidgetNotes
texttextarea
numberinput
booleancheckbox
ordinalselectRequires options=[…].
ds.add_column("verdict", "ordinal", options=["A", "B", "C", "D"], required=True)
ds.columns() # [Column(id=…, name='input', position=0, type='text', required=False), …]

Existing datasets come back by id or all at once:

lab.datasets() # [Dataset, …] in this workspace
lab.dataset("8f21…") # one you already have

Two verbs, and the difference is not convenience.

done = ds.upload(rows, message="MedQA test, first 500")
done.op # 'import'
done.n # 1 — this dataset's first version
done.rows # 500 — how many THIS upload put in
done.total # 500 — how many the table has now

upload goes through the platform’s two-phase import, so it names what it wrote: the version records the verb, the message, and where the rows came from.

ds.add(rows) # writes rows, mints nothing. Returns the new total.
ds.add_returning(rows) # the same, but hands back the written rows with their ids

add is the short path, and honest about what it does not do: it leaves the table with work no version points at, and that is what a later filter, sample, or experiment pinned to a version refuses to stand on.

ds.upload(more, message="another 200", mode="append") # default
ds.upload(everything, message="rebuilt", mode="replace")

replace changes the whole table — and re-mints the identity of every row. The diff against the previous version will say they all left and as many arrived, even where the content is identical. To add without losing identity, append.

ds.upload_csv(open("medqa.csv").read(), message="the file as it came", name="medqa.csv")

name is the file name: it goes into the version’s provenance, so the lineage says what it was imported from rather than only that it was imported.

ds.count() # 500, without fetching anything
ds.rows() # one page — 1000 at most
ds.rows(limit=50, offset=100) # one page, where you asked
ds.all_rows() # every page, walked for you

Each row is its data with its id folded in:

ds.rows()[0]
# {'input': 'A 45-year-old man…', 'reference': 'C', 'id': '7c1e…'}

The server serves Dataset.PAGE_MAX — 1000 — rows at most per page. Asking for more raises rather than silently trimming; all_rows() is the loop, written once, instead of one off-by-one per caller.

A dataset with a column literally called id cannot be flattened this way without losing either the column or the identity, so it is refused out loud. Read that one with Platform.call.

Platform.call is public on purpose. This wheel covers datasets; the platform is much larger, and the alternative is a second urllib written beside it.

graph = lab.call("POST", "/graphs", {"name": "medqa panel", "graph_data": blob})["id"]
experiment = lab.call("POST", "/experiments", {
"graph_id": graph,
"name": "panel v2",
"dataset_id": ds.id,
"param_space": {
"type": "grid",
"params": [],
"score": {"metric": "exact_match", "direction": "maximize"},
},
})["id"]
lab.call("POST", f"/experiments/{experiment}/pin", {"graph_version_id": version_id})
lab.call("POST", f"/experiments/{experiment}/run", {})
lab.call("GET", f"/experiments/{experiment}/trials")

The bearer, the workspace header and the error handling are already done. A refusal raises PlatformError, whose message is the body the platform sent — which is where it says what was actually wrong — and whose .status is the HTTP code when there was one.

try:
ds.upload(rows)
except cl.PlatformError as no:
no.status # 400, 401, 404… or None if the call never landed
ds.delete()