Skip to main content
An RL run trains one policy for one robot in simulation with reinforcement learning. You send a task spec: the robot, what to learn, and how long to train. Kite checks the spec for free, trains it on a GPU, then packages a bundle. The bundle holds the policy as ONNX and a checkpoint, the MuJoCo scene it trained in with a script that runs it there, clips, and a report whose verdict comes from measured checks. Every request and response on this page comes from one real run. It fine-tunes the published Open Duck Mini v2 walk for 500 iterations on an L4.

See it in action

The run was given the published walk as its starting point and this description, which the report’s judge later compares the clip against:
Description”Walk forward at a steady pace with a level body and alternating feet.”
Published walk
After this run

Both policies commanded to walk at 0.3 m/s, rendered from the same start with the same camera. 500 iterations of fine-tuning (about 7 minutes of training on an L4) raised velocity tracking from 0.598 to 0.648 and brought the torso from 4.3° to 3.6° off level, measured the same way for both. Run outside Kite, the fine-tuned policy covers 2.94 m in 10 seconds without falling.

Check a spec first

POST /v1/rl_runs/validate takes the same body as a create. It compiles the spec and rolls three canned policies through it on a CPU: zero action, random actions, and, for robots that have one, a published walk that passes the report’s checks. It reports what each reward term pays each of them, and warns when a reward pays for the wrong thing. It creates nothing and costs nothing.
resolved_params lists every parameter as the run would train with it (all 48; a few are shown here). The rollouts are the scale for everything that follows. A trained walk earns about 7 per step here, standing still 5.7, and flailing 3.4 with a fall every two seconds. When standing still earns most of the tracking reward, validate returns idle_policy_scores_high with a patch: the change to your spec that fixes it. POST /v1/rl_runs/estimate returns only the estimate. Tokens are reserved when you create the run, and whatever training doesn’t use is refunded.

Create an RL run

This is the request the example sent, walk.json:
walk.json
It returns 202 with one run per seed, all sharing a group_id (abridged; the full object is at the end of this page):
Only robot and task.objective are required. Everything else has a verified default, and the run object echoes the resolved spec back.

Request body

string
required
open_duck_mini_v2, microduck, or unitree_g1. GET /v1/rl_runs/catalog lists each robot with its command band and measured training throughput.
object
required
What to learn. objective is velocity (walk, turn, or stand on command), imitation (follow a reference motion, motion), or balance (stand on a scene item, item). commands gives velocity runs the ranges each episode’s command is drawn from; any range you leave out keeps the robot’s verified band. scene.items adds physical items (a sphere, box, or capsule with a size, position, and mass). params overrides any reward or randomization parameter; the catalog lists every one with its default and bounds.
object
Start from another policy instead of from scratch: a public policy ({"policy": "4d574ce4c547"}) or one of your runs ({"run": "rlr_…"}). mode is fork (its weights and training schedule, with a fresh optimizer) or continue (also its optimizer; your runs only). checkpoint picks a saved iteration of a run.
object
preset is probe (300 iterations, a first signal in minutes) or standard (the iterations each objective was verified with: 4,000 for velocity). iterations overrides the preset. num_envs sets the parallel worlds (default 2,048). max_tokens stops training, keeping its policy, once the run has used that many tokens.
integer | integer[]
A count (1–5) or explicit seeds. One run per seed, sharing a group_id.
object
Extras beyond what every run produces. formats adds torchscript. videos adds a clip per command, up to four. huggingface also uploads the bundle to your connected Hugging Face account.
string
What the behavior should look like, in plain words. The report’s judge compares the clip against it.
string
The dated trainer, for example 2026-09-28. Defaults to the latest, or to the parent’s for a fork or continue of your own run. A run keeps its engine, so a spec trains the same way until you change the date.
hardware_tier, visibility, title, metadata, and webhook_metadata are described in the API reference.

Idempotency

Send an Idempotency-Key to make a retried create safe. A retry with the same key and body returns the same runs, even while the first request is still in flight. The same key with a different body returns 409 idempotency_key_reused.

Track progress

Poll GET /v1/rl_runs/:id, GET /v1/operations/:id, or subscribe to rl_run.completed, rl_run.failed, and rl_run.canceled webhooks. A run moves queued → processing → succeeded. Within processing, phase says whether it is starting, training, or packaging. A canceled run that has trained shows canceling while its policy is packaged. The example’s status_message read:
Time spent waiting for a machine isn’t billed. While it trains, the run’s metrics carry the latest iteration, the mean reward per step, the same broken down by reward term, and throughput. GET /v1/rl_runs/:id/metrics returns the whole curve, every 5 iterations:
Mean reward per step over 500 iterations, and the reward per step of each term: track, upright, pose, clearance, feet, slip, and regularizers

The example's training curve, from GET /v1/rl_runs/:id/metrics. Tracking rises while the regularizers term grows more negative as the smoothness penalties ramp up.

The report

When training ends, Kite compares the last saved checkpoints and keeps the best. It then measures that checkpoint over 24 worlds for 12 seconds each with training noise on, renders it, and exports it. GET /v1/rl_runs/:id/report returns the verdict and how it was reached: The verdict is pass when every check passes, and fail otherwise. The judge, a vision model, reads 8 frames of the clean clip against your description:
The judge advises; it doesn’t certify. It can turn a passing run into needs_review when it disagrees with high confidence, but it can never pass a run that failed a check. Its suggestions are reward parameters to scale in a fork. The bundle’s contact sheet shows the same clip:
Sixteen frames of Open Duck Mini v2 walking upright with alternating steps

media/sheet.jpg from the example's bundle: 16 frames of the clean clip, the policy walking at 0.3 m/s.

Download the bundle

Once the run is packaged, its output says where everything is. GET /v1/rl_runs/:id/archive returns the bundle as one zip that unpacks to <run id>/; check it against output.sha256. GET /v1/rl_runs/:id/files lists every file with its size and sha256, and GET /v1/rl_runs/:id/files/{path} returns one.
The example’s bundle is 12.3 MB:
policy/policy.json describes the policy so you can run it anywhere: nine observation blocks (angular velocity, gravity direction, joint positions and velocities, the previous action, the command, foot contact, height, and air time), the action rule (default_pose + scale × action, clipped to the joint limits), PD gains, and the control rate. Any saved checkpoint is also available on its own from GET /v1/rl_runs/:id/checkpoints/:iteration/download, as ONNX by default or with ?format=pt.

Run it

play.py needs only MuJoCo and onnxruntime:
python play.py opens the MuJoCo viewer (on macOS, mjpython play.py), --vx, --vy, and --yaw change the command, and --headless --video out.mp4 records a clip. This is the clip the example asked for in outputs.videos, 8 seconds at 0.3 m/s:

media/cmd_vx0.3.mp4 from the example's bundle.

Iterate

Most runs aren’t right the first time, so the API is built for short loops:
  • Probe first. "budget": {"preset": "probe"} trains 300 iterations: enough to see whether the reward teaches the right thing, in minutes. Scale up the spec that works.
  • Fork what almost works. "from": {"run": "rlr_…"} starts from a run’s weights and training schedule with this spec, so changing one reward parameter doesn’t mean training from scratch. The judge’s suggestions are a good first change. "mode": "continue" also keeps its optimizer, to simply train longer.
  • Run seeds side by side. "seeds": 3 starts three runs of one spec in one group. RL results vary by seed, and the SDK’s best() picks the winner.
  • Cap the spend. budget.max_tokens stops a run at that many tokens and keeps what it learned.
A probe is a first signal, not a finished policy. Here is a 300-iteration probe of the same walk, trained from scratch, next to the example’s fork of the published walk:
Probe: 300 iterations
Fork: 500 iterations

Left, a probe trained from scratch; right, the example’s fork of the published walk. The probe already walks: tracking 0.569, no falls in 10 seconds. It passes every check, but leans 10.4° forward, and its judge noted the lean and mushy foot phasing. The fork starts from a walk that is already good and refines it.

Without writing HTTP

The Kite CLI and Python SDK wrap the whole flow: validate, train, wait, download, and verify.
kite rl fork rlr_… --set air_time=4.2 forks a run with one parameter changed; that is the judge’s ×1.4 on the default of 3.0. Agents get the same flow from the MCP tools kite_rl_validate, kite_rl_train, and kite_rl_status, whose next field is the exact download command once the run is packaged.

Cancel

POST /v1/rl_runs/:id/cancel stops a run. A run that has started training keeps the policy it has so far: it reads canceling while that policy is packaged, then canceled, with its outputs, and is charged for the minutes it trained. A run that never started, or left nothing to package, is refunded in full.

Billing

Training is billed by the minute at the tier’s rate. On gcp_gpu_l4 that is 1,000 tokens per hour. Tokens are reserved at create (the estimate, or budget.max_tokens if lower) and settled when the run ends. The example reserved 167 tokens and was charged 117, for 7 minutes of training; 50 came back. You’re charged only for a run that produces a policy. A failed run, and time spent waiting for a machine, cost nothing. Usage appears under rl_runs in GET /v1/usage.

Errors

A spec is refused before anything is reserved, with a code and the field to fix in param:
  • 400 parameter_invalid: an unknown robot, a command range outside the robot’s limits, or a scene item out of bounds.
  • 400 unknown_param or param_out_of_range: a task.params key the objective doesn’t have, or a value outside its bounds. The catalog lists both.
  • 400 parameter_required: for example, an imitation run without task.motion.
  • 402 insufficient_tokens: details has tokens_required and current_balance.
  • 429 concurrency_limit_exceeded: too many runs in progress. Wait for one to finish, or cancel it.
  • 409 run_not_ready: the report, files, or archive of a run that isn’t packaged yet.
A run that fails carries an error with a code: warm_start_failed, training_diverged, gpu_capacity_unavailable, training_failed, or packaging_failed. Nothing is charged for any of them.

The RL run object

The example, once it finished:
Every field is documented in the API reference, generated from the API itself.