> ## Documentation Index
> Fetch the complete documentation index at: https://docs.kiteml.com/llms.txt
> Use this file to discover all available pages before exploring further.

# RL runs: train a robot policy in simulation from a task spec

> Describe the behavior, robot, and budget as JSON. Kite trains the policy with reinforcement learning on a GPU and returns it as ONNX, with the MuJoCo scene it trained in, clips, and a verdict from measured checks.

An **RL run** trains one policy for one robot in simulation with reinforcement learning. You send a task spec: the robot, what to learn, and how long to train. Kite checks the spec for free, trains it on a GPU, then packages a bundle. The bundle holds the policy as ONNX and a checkpoint, the MuJoCo scene it trained in with a script that runs it there, clips, and a report whose verdict comes from measured checks.

Every request and response on this page comes from one real run. It fine-tunes the published Open Duck Mini v2 walk for 500 iterations on an L4.

## See it in action

The run was given the published walk as its starting point and this description, which the report's judge later compares the clip against:

<div className="kite-prompt">
  <span className="kite-prompt__label">Description</span>
  <span className="kite-prompt__text">"Walk forward at a steady pace with a level body and alternating feet."</span>
</div>

<div className="kite-vs">
  <div className="kite-vs__panel">
    <span className="kite-vs__label">Published walk</span>

    <video src="https://mintcdn.com/kite-ml/kRWAN4UkuvurYqj2/videos/rl-runs/published-walk.mp4?fit=max&auto=format&n=kRWAN4UkuvurYqj2&q=85&s=675b2a4813fdb644e3bf042ea361593a" autoPlay loop muted playsInline preload="metadata" data-path="videos/rl-runs/published-walk.mp4" />
  </div>

  <div className="kite-vs__panel">
    <span className="kite-vs__label kite-vs__label--after">After this run</span>

    <video src="https://mintcdn.com/kite-ml/kRWAN4UkuvurYqj2/videos/rl-runs/fine-tune.mp4?fit=max&auto=format&n=kRWAN4UkuvurYqj2&q=85&s=098ae99dd4e77978331f773fa168199d" autoPlay loop muted playsInline preload="metadata" data-path="videos/rl-runs/fine-tune.mp4" />
  </div>
</div>

<p className="kite-vs__caption">Both policies commanded to walk at 0.3 m/s, rendered from the same start with the same camera. 500 iterations of fine-tuning (about 7 minutes of training on an L4) raised velocity tracking from 0.598 to 0.648 and brought the torso from 4.3° to 3.6° off level, measured the same way for both. Run outside Kite, the fine-tuned policy covers 2.94 m in 10 seconds without falling.</p>

## Check a spec first

`POST /v1/rl_runs/validate` takes the same body as a create. It compiles the spec and rolls three canned policies through it on a CPU: zero action, random actions, and, for robots that have one, a published walk that passes the report's checks. It reports what each reward term pays each of them, and warns when a reward pays for the wrong thing. It creates nothing and costs nothing.

```bash theme={"system"}
curl https://api.kiteml.com/v1/rl_runs/validate \
  -H "Authorization: Bearer $KITE_API_KEY" \
  -H "Kite-Version: 2026-09-27" \
  -H "Content-Type: application/json" \
  -d @walk.json
```

```json theme={"system"}
{
  "object": "rl_validation",
  "ok": true,
  "observation_size": 57,
  "action_size": 14,
  "control_hz": 50.0,
  "rollouts": [
    { "policy": "zero_action", "reward_per_step": 5.7444, "falls_per_min": 0.0, "nonfinite": 0,
      "terms": { "track": 2.7529, "upright": 2.0, "pose": 0.9923, "head": 0.0, "feet": -0.0006, "clearance": 0.0003,
                 "slip": -0.0, "regularizers": -0.0005 } },
    { "policy": "random", "reward_per_step": 3.3678, "falls_per_min": 33.75, "nonfinite": 0,
      "terms": { "track": 1.6334, "upright": 1.808, "pose": 0.8588, "head": 0.0, "feet": -0.0066, "clearance": 0.143,
                 "slip": -0.0024, "regularizers": -1.0664 } },
    { "policy": "catalog:4d574ce4c547", "reward_per_step": 6.9638, "falls_per_min": 0.0, "nonfinite": 0,
      "terms": { "track": 3.7559, "upright": 1.9984, "pose": 0.8585, "head": 0.0, "feet": 0.0343, "clearance": 0.3969,
                 "slip": -0.0034, "regularizers": -0.0768 } }
  ],
  "warnings": [],
  "resolved_params": { "track_lin": 2.0, "sigma_lin": 0.1, "track_yaw": 2.0, "pose": 1.0, "head_pose": 0.0,
                       "air_time": 3.0, "clearance": 1.0, "action_rate_final": 1.0 },
  "estimate": { "object": "rl_estimate", "minutes_per_run": 9.1, "tokens_per_run": 167, "runs": 1, "tokens": 167,
                "capped": false, "startup_minutes": 5 }
}
```

`resolved_params` lists every parameter as the run would train with it (all 48; a few are shown here). The rollouts are the scale for everything that follows. A trained walk earns about 7 per step here, standing still 5.7, and flailing 3.4 with a fall every two seconds. When standing still earns most of the tracking reward, validate returns `idle_policy_scores_high` with a `patch`: the change to your spec that fixes it.

`POST /v1/rl_runs/estimate` returns only the estimate. Tokens are reserved when you create the run, and whatever training doesn't use is refunded.

## Create an RL run

This is the request the example sent, `walk.json`:

```json walk.json theme={"system"}
{
  "robot": "open_duck_mini_v2",
  "title": "Open Duck: steady walk (fine-tuned)",
  "description": "Walk forward at a steady pace with a level body and alternating feet.",
  "task": {
    "objective": "velocity",
    "commands": { "vx": [0.2, 0.4], "vy": [-0.05, 0.05], "yaw": [-0.2, 0.2] }
  },
  "from": { "policy": "4d574ce4c547" },
  "budget": { "iterations": 500, "max_tokens": 400 },
  "seeds": 1,
  "hardware_tier": "gcp_gpu_l4",
  "outputs": {
    "formats": ["onnx", "pt", "torchscript"],
    "videos": [{ "command": { "vx": 0.3 }, "seconds": 8 }]
  },
  "metadata": { "example": "rl_runs/walk.json" }
}
```

```bash theme={"system"}
curl https://api.kiteml.com/v1/rl_runs \
  -H "Authorization: Bearer $KITE_API_KEY" \
  -H "Kite-Version: 2026-09-27" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d @walk.json
```

It returns `202` with one run per seed, all sharing a `group_id` (abridged; the full object is [at the end of this page](#the-rl-run-object)):

```json theme={"system"}
{
  "object": "list",
  "data": [
    {
      "id": "rlr_01M3MQ39HVAEBQ3F64K6JDF9XR",
      "object": "rl_run",
      "status": "queued",
      "status_message": "waiting for a machine",
      "progress": 0.0,
      "robot": "open_duck_mini_v2",
      "engine": "2026-09-28",
      "group_id": "rlg_01M3MQ39HV4YNTWNNY2GHM55VC",
      "seed": 0,
      "from": { "policy": "4d574ce4c547", "mode": "fork" },
      "hardware_tier": "gcp_gpu_l4",
      "tokens": { "reserved": 167, "charged": 0, "refunded": 0 },
      "created_at": "2026-09-28T19:15:36.753577Z"
    }
  ],
  "has_more": false,
  "next_cursor": null
}
```

Only `robot` and `task.objective` are required. Everything else has a verified default, and the run object echoes the resolved spec back.

### Request body

<ParamField body="robot" type="string" required>
  `open_duck_mini_v2`, `microduck`, or `unitree_g1`. `GET /v1/rl_runs/catalog` lists each robot with its command band and measured training throughput.
</ParamField>

<ParamField body="task" type="object" required>
  What to learn. `objective` is `velocity` (walk, turn, or stand on command), `imitation` (follow a reference motion, `motion`), or `balance` (stand on a scene item, `item`). `commands` gives velocity runs the ranges each episode's command is drawn from; any range you leave out keeps the robot's verified band. `scene.items` adds physical items (a sphere, box, or capsule with a size, position, and mass). `params` overrides any reward or randomization parameter; the catalog lists every one with its default and bounds.
</ParamField>

<ParamField body="from" type="object">
  Start from another policy instead of from scratch: a public policy (`{"policy": "4d574ce4c547"}`) or one of your runs (`{"run": "rlr_…"}`). `mode` is `fork` (its weights and training schedule, with a fresh optimizer) or `continue` (also its optimizer; your runs only). `checkpoint` picks a saved iteration of a run.
</ParamField>

<ParamField body="budget" type="object">
  `preset` is `probe` (300 iterations, a first signal in minutes) or `standard` (the iterations each objective was verified with: 4,000 for velocity). `iterations` overrides the preset. `num_envs` sets the parallel worlds (default 2,048). `max_tokens` stops training, keeping its policy, once the run has used that many tokens.
</ParamField>

<ParamField body="seeds" type="integer | integer[]">
  A count (1–5) or explicit seeds. One run per seed, sharing a `group_id`.
</ParamField>

<ParamField body="outputs" type="object">
  Extras beyond what every run produces. `formats` adds `torchscript`. `videos` adds a clip per command, up to four. `huggingface` also uploads the bundle to your connected Hugging Face account.
</ParamField>

<ParamField body="description" type="string">
  What the behavior should look like, in plain words. The report's judge compares the clip against it.
</ParamField>

<ParamField body="engine" type="string">
  The dated trainer, for example `2026-09-28`. Defaults to the latest, or to the parent's for a fork or continue of your own run. A run keeps its engine, so a spec trains the same way until you change the date.
</ParamField>

`hardware_tier`, `visibility`, `title`, `metadata`, and `webhook_metadata` are described in the [API reference](/platform-api/api-reference).

### Idempotency

Send an `Idempotency-Key` to make a retried create safe. A retry with the same key and body returns the same runs, even while the first request is still in flight. The same key with a different body returns `409 idempotency_key_reused`.

## Track progress

Poll `GET /v1/rl_runs/:id`, `GET /v1/operations/:id`, or subscribe to `rl_run.completed`, `rl_run.failed`, and `rl_run.canceled` webhooks. A run moves `queued → processing → succeeded`. Within `processing`, `phase` says whether it is `starting`, `training`, or `packaging`. A canceled run that has trained shows `canceling` while its policy is packaged. The example's `status_message` read:

```
queued: waiting for a machine
starting: compiling the scene
training: iteration 250 of 500
packaging: choosing a checkpoint, measuring, rendering and exporting
done: verdict pass
```

Time spent waiting for a machine isn't billed. While it trains, the run's `metrics` carry the latest iteration, the mean reward per step, the same broken down by reward term, and throughput. `GET /v1/rl_runs/:id/metrics` returns the whole curve, every 5 iterations:

```json theme={"system"}
{
  "iteration": 500,
  "mean_reward": 6.39074,
  "terms": {
    "track": 3.61929, "upright": 1.99652, "pose": 0.81887, "head": 0.0,
    "clearance": 0.46359, "feet": 0.04187, "slip": -0.00441, "regularizers": -0.545
  },
  "episode_s": 162.63,
  "steps_per_second": 59745,
  "elapsed": 411.4
}
```

<Frame caption="The example's training curve, from GET /v1/rl_runs/:id/metrics. Tracking rises while the regularizers term grows more negative as the smoothness penalties ramp up.">
  <img src="https://mintcdn.com/kite-ml/kRWAN4UkuvurYqj2/images/rl-runs/training-curve.png?fit=max&auto=format&n=kRWAN4UkuvurYqj2&q=85&s=73c8c45bac298827baf1f81a0bc87132" alt="Mean reward per step over 500 iterations, and the reward per step of each term: track, upright, pose, clearance, feet, slip, and regularizers" width="1760" height="624" data-path="images/rl-runs/training-curve.png" />
</Frame>

## The report

When training ends, Kite compares the last saved checkpoints and keeps the best. It then measures that checkpoint over 24 worlds for 12 seconds each with training noise on, renders it, and exports it. `GET /v1/rl_runs/:id/report` returns the verdict and how it was reached:

| Check | Example | Limit |
| - | - | - |
| `falls_per_min` | 0.42 | ≤ 2 |
| `survival` | 0.917 | ≥ 0.75 |
| `pitch_offset_deg` | 3.55 | ≤ 15 |
| `tracking` | 0.648 | ≥ 0.5 |
| `alternation` | 0.455 | ≥ 0.2 |

The verdict is `pass` when every check passes, and `fail` otherwise. The judge, a vision model, reads 8 frames of the clean clip against your `description`:

```json theme={"system"}
{
  "matches": true,
  "confidence": 0.72,
  "diagnosis": "The frames show the robot translating steadily toward the camera with a level torso and consistent upright posture, with feet swinging past one another rather than shuffling in place, so the forward-walk command is met. Gait quality is mediocre though: alternation is only 0.455 (partially in-phase / irregular swing timing) and lin-vel tracking is 0.648 with a slight yaw drift, so the pace is not perfectly steady or straight.",
  "suggestions": [
    { "weight": "air_time", "factor": 1.4 },
    { "weight": "track_lin", "factor": 1.2 },
    { "weight": "slip", "factor": 1.2 }
  ]
}
```

The judge advises; it doesn't certify. It can turn a passing run into `needs_review` when it disagrees with high confidence, but it can never pass a run that failed a check. Its `suggestions` are reward parameters to scale in a fork. The bundle's contact sheet shows the same clip:

<Frame caption="media/sheet.jpg from the example's bundle: 16 frames of the clean clip, the policy walking at 0.3 m/s.">
  <img src="https://mintcdn.com/kite-ml/kRWAN4UkuvurYqj2/images/rl-runs/contact-sheet.jpg?fit=max&auto=format&n=kRWAN4UkuvurYqj2&q=85&s=173ff75fce0953b68e08f870da6c9979" alt="Sixteen frames of Open Duck Mini v2 walking upright with alternating steps" width="1440" height="810" data-path="images/rl-runs/contact-sheet.jpg" />
</Frame>

## Download the bundle

Once the run is packaged, its `output` says where everything is. `output.video_url` is the policy on video and `output.policy_url` the trained policy, one request each. `GET /v1/rl_runs/:id/archive` returns the whole bundle as one zip, `kiteml_<run id>.zip`, that unpacks to `kiteml_<run id>/`; check it against `output.sha256`. `GET /v1/rl_runs/:id/files` lists every file with its size and sha256, and `GET /v1/rl_runs/:id/files/{path}` returns one.

```bash theme={"system"}
curl -fL https://api.kiteml.com/v1/rl_runs/rlr_01M3MQ39HVAEBQ3F64K6JDF9XR/archive \
  -H "Authorization: Bearer $KITE_API_KEY" \
  -H "Kite-Version: 2026-09-27" \
  -o kiteml_rlr_01M3MQ39HVAEBQ3F64K6JDF9XR.zip
```

The top of the bundle holds the README, the report, and one folder per kind of output. The example's bundle is 12.3 MB:

```
kiteml_rlr_01M3MQ39HVAEBQ3F64K6JDF9XR/
  README.md                what each folder holds, the verdict, how to run it
  report.json              the verdict, its checks, the judge
  media/video.mp4          a clean clip of the policy
  media/video_vx0.3.mp4    the clip outputs.videos asked for
  media/sheet.jpg          16 frames of the clean clip (thumbnail.jpg is its middle frame)
  policy/policy.onnx       the policy: 57 raw observations in, 14 joint targets out, 50 Hz (normalizer built in)
  policy/policy.json       the contract: every observation block, joint order, gains, rates
  policy/policy.jit.pt     TorchScript (outputs.formats)
  mujoco/scene.xml         the MuJoCo model the policy trained in
  mujoco/assets/           its meshes
  mujoco/obs.py            builds the observation, block by block
  mujoco/play.py           runs the policy in MuJoCo, with no Kite code
  training/policy.pt       the training checkpoint, for fine-tuning
  training/metrics.csv     the training curve, by reward term
  training/spec.json       everything needed to train it again
```

`policy/policy.json` describes the policy so you can run it anywhere: nine observation blocks (angular velocity, gravity direction, joint positions and velocities, the previous action, the command, foot contact, height, and air time), the action rule (`default_pose + scale × action`, clipped to the joint limits), PD gains, and the control rate. Any saved checkpoint is also available on its own from `GET /v1/rl_runs/:id/checkpoints/:iteration/download`, as ONNX by default or with `?format=pt`.

### Run it

`play.py` needs only MuJoCo and onnxruntime:

```bash theme={"system"}
pip install mujoco onnxruntime numpy
cd kiteml_rlr_01M3MQ39HVAEBQ3F64K6JDF9XR/mujoco
python play.py --headless --seconds 10
```

```json theme={"system"}
{"seconds": 10.0, "distance_m": 2.939, "falls": 0, "command": [0.3, 0.0, 0.0], "video": null}
```

`python play.py` opens the MuJoCo viewer (on macOS, `mjpython play.py`), `--vx`, `--vy`, and `--yaw` change the command, and `--headless --video out.mp4` records a clip. This is the clip the example asked for in `outputs.videos`, 8 seconds at 0.3 m/s:

<Frame caption="media/video_vx0.3.mp4 from the example's bundle.">
  <video src="https://mintcdn.com/kite-ml/kRWAN4UkuvurYqj2/videos/rl-runs/fine-tune-0.3ms.mp4?fit=max&auto=format&n=kRWAN4UkuvurYqj2&q=85&s=86939410e468bcd7364ae674367ff727" autoPlay loop muted playsInline preload="metadata" data-path="videos/rl-runs/fine-tune-0.3ms.mp4" />
</Frame>

## Iterate

Most runs aren't right the first time, so the API is built for short loops:

* **Probe first.** `"budget": {"preset": "probe"}` trains 300 iterations: enough to see whether the reward teaches the right thing, in minutes. Scale up the spec that works.
* **Fork what almost works.** `"from": {"run": "rlr_…"}` starts from a run's weights and training schedule with this spec, so changing one reward parameter doesn't mean training from scratch. The judge's `suggestions` are a good first change. `"mode": "continue"` also keeps its optimizer, to simply train longer.
* **Run seeds side by side.** `"seeds": 3` starts three runs of one spec in one group. RL results vary by seed, and the SDK's `best()` picks the winner.
* **Cap the spend.** `budget.max_tokens` stops a run at that many tokens and keeps what it learned.

A probe is a first signal, not a finished policy. Here is a 300-iteration probe of the same walk, trained from scratch, next to the example's fork of the published walk:

<div className="kite-vs">
  <div className="kite-vs__panel">
    <span className="kite-vs__label">Probe: 300 iterations</span>

    <video src="https://mintcdn.com/kite-ml/kRWAN4UkuvurYqj2/videos/rl-runs/probe.mp4?fit=max&auto=format&n=kRWAN4UkuvurYqj2&q=85&s=1b9dcf1de14551563d85b92fead5b7a3" autoPlay loop muted playsInline preload="metadata" data-path="videos/rl-runs/probe.mp4" />
  </div>

  <div className="kite-vs__panel">
    <span className="kite-vs__label kite-vs__label--after">Fork: 500 iterations</span>

    <video src="https://mintcdn.com/kite-ml/kRWAN4UkuvurYqj2/videos/rl-runs/fine-tune.mp4?fit=max&auto=format&n=kRWAN4UkuvurYqj2&q=85&s=098ae99dd4e77978331f773fa168199d" autoPlay loop muted playsInline preload="metadata" data-path="videos/rl-runs/fine-tune.mp4" />
  </div>
</div>

<p className="kite-vs__caption">Left, a probe trained from scratch; right, the example's fork of the published walk. The probe already walks: tracking 0.569, no falls in 10 seconds. It passes every check, but leans 10.4° forward, and its judge noted the lean and mushy foot phasing. The fork starts from a walk that is already good and refines it.</p>

## Without writing HTTP

The Kite CLI and Python SDK wrap the whole flow: validate, train, wait, download, and verify.

<CodeGroup>
  ```bash CLI theme={"system"}
  kite rl validate walk.json
  kite rl train walk.json --wait
  kite rl download rlr_01M3MQ39HVAEBQ3F64K6JDF9XR -o ./policies
  ```

  ```python Python theme={"system"}
  import json
  from kite_sdk import Kite

  kite = Kite()
  spec = json.load(open("walk.json"))
  print(kite.rl_runs.validate(**spec)["warnings"])
  runs = kite.rl_runs.create(**spec)                  # one run per seed
  best = kite.rl_runs.best(r.wait(on_update=print) for r in runs)
  path = best.download("./policies")                  # checked against output.sha256
  ```
</CodeGroup>

`kite rl fork rlr_… --set air_time=4.2` forks a run with one parameter changed; that is the judge's ×1.4 on the default of 3.0. Agents get the same flow from the MCP tools `kite_rl_validate`, `kite_rl_train`, and `kite_rl_status`, whose `next` field is the exact download command once the run is packaged.

## Cancel

`POST /v1/rl_runs/:id/cancel` stops a run. A run that has started training keeps the policy it has so far: it reads `canceling` while that policy is packaged, then `canceled`, with its outputs, and is charged for the minutes it trained. A run that never started, or left nothing to package, is refunded in full.

## Billing

Training is billed by the minute at the tier's rate. On `gcp_gpu_l4` that is 1,000 tokens per hour. Tokens are reserved at create (the estimate, or `budget.max_tokens` if lower) and settled when the run ends. The example reserved 167 tokens and was charged 117, for 7 minutes of training; 50 came back. You're charged only for a run that produces a policy. A failed run, and time spent waiting for a machine, cost nothing. Usage appears under `rl_runs` in `GET /v1/usage`.

## Errors

A spec is refused before anything is reserved, with a code and the field to fix in `param`:

* `400 parameter_invalid`: an unknown robot, a command range outside the robot's limits, or a scene item out of bounds.
* `400 unknown_param` or `param_out_of_range`: a `task.params` key the objective doesn't have, or a value outside its bounds. The catalog lists both.
* `400 parameter_required`: for example, an imitation run without `task.motion`.
* `402 insufficient_tokens`: `details` has `tokens_required` and `current_balance`.
* `429 concurrency_limit_exceeded`: too many runs in progress. Wait for one to finish, or cancel it.
* `409 run_not_ready`: the report, files, or archive of a run that isn't packaged yet.

A run that fails carries an `error` with a `code`: `warm_start_failed`, `training_diverged`, `gpu_capacity_unavailable`, `training_failed`, or `packaging_failed`. Nothing is charged for any of them.

## The RL run object

The example, once it finished:

```json theme={"system"}
{
  "id": "rlr_01M3MQ39HVAEBQ3F64K6JDF9XR",
  "object": "rl_run",
  "status": "succeeded",
  "phase": null,
  "status_message": "done: verdict pass",
  "progress": 1.0,
  "robot": "open_duck_mini_v2",
  "engine": "2026-09-28",
  "group_id": "rlg_01M3MQ39HV4YNTWNNY2GHM55VC",
  "seed": 0,
  "title": "Open Duck: steady walk (fine-tuned)",
  "description": "Walk forward at a steady pace with a level body and alternating feet.",
  "task": {
    "objective": "velocity",
    "commands": { "vx": [0.2, 0.4], "vy": [-0.05, 0.05], "yaw": [-0.2, 0.2] },
    "motion": null,
    "item": null,
    "scene": { "items": [] },
    "params": {}
  },
  "budget": { "preset": null, "iterations": 500, "num_envs": 2048, "max_tokens": 400 },
  "from": { "policy": "4d574ce4c547", "mode": "fork" },
  "outputs": {
    "formats": ["onnx", "pt", "torchscript"],
    "videos": [{ "command": { "vx": 0.3 }, "seconds": 8.0 }],
    "huggingface": null
  },
  "hardware_tier": "gcp_gpu_l4",
  "visibility": "private",
  "metrics": {
    "iteration": 500,
    "max_iterations": 500,
    "env_steps": 24576000,
    "mean_reward": 6.39074,
    "steps_per_second": 59745,
    "elapsed_s": 411.4,
    "terms": { "track": 3.61929, "upright": 1.99652, "pose": 0.81887, "head": 0.0, "clearance": 0.46359,
               "feet": 0.04187, "slip": -0.00441, "regularizers": -0.545 }
  },
  "report": { "verdict": "pass", "best_checkpoint": 500 },
  "output": {
    "policy_url": "/v1/rl_runs/rlr_01M3MQ39HVAEBQ3F64K6JDF9XR/files/policy/policy.onnx",
    "video_url": "/v1/rl_runs/rlr_01M3MQ39HVAEBQ3F64K6JDF9XR/files/media/video.mp4",
    "files": 43,
    "bytes": 12294261,
    "sha256": "d085065b6d2bf18d5b0bffd67c8b8ed0834136e446d608fce13834ea2976706f",
    "archive_url": "/v1/rl_runs/rlr_01M3MQ39HVAEBQ3F64K6JDF9XR/archive",
    "files_url": "/v1/rl_runs/rlr_01M3MQ39HVAEBQ3F64K6JDF9XR/files",
    "report_url": "/v1/rl_runs/rlr_01M3MQ39HVAEBQ3F64K6JDF9XR/report",
    "huggingface": null
  },
  "error": null,
  "tokens": { "reserved": 167, "charged": 117, "refunded": 50 },
  "metadata": { "example": "rl_runs/walk.json" },
  "webhook_metadata": null,
  "created_at": "2026-09-28T19:15:36.753577Z",
  "started_at": "2026-09-28T19:51:13.239324Z",
  "completed_at": "2026-09-28T19:59:53.818362Z"
}
```

Every field is documented in the [API reference](/platform-api/api-reference), generated from the API itself.
