Full project docs: README.md, SPEC.md, architecture, setup guide
The web UI of Corral serves a REST API at port 8006 (corral web), plus
WebSocket bridges for VNC and serial consoles. All responses are JSON unless
noted otherwise.
The same API is available from the on-cluster deployment at
https://corral.<tailnet>.ts.net.
Every error response has shape {"error": "<message>"} with a 4xx or 5xx
status code. 5xx means the server couldn't reach the cluster (kubectl failed);
4xx means invalid input.
List all VMs across all namespaces. Merges live VMI data (IP, node) for VMs that run.
Response (array):
[
{
"name": "web",
"namespace": "tailvm",
"status": "Running",
"ready": true,
"running": true,
"cpu": 2,
"mem": "4Gi",
"disk": "20Gi",
"ip": "10.42.0.15",
"node": "bihar"
}
]Create a VM.
Request fields:
| Field | Type | Required | Notes |
|---|---|---|---|
name |
string | yes | [a-z0-9][a-z0-9-]* |
namespace |
string | no | defaults to tailvm |
cpu |
int | no | default 2 |
mem |
string | no | e.g. 4G, default 4G |
disk |
string | no | e.g. 20G, default 20G |
containerDisk |
string | source | OCI containerdisk URI |
image |
string | source | catalog entry name (see /api/images) |
import |
string | source | qcow2/raw URL (CDI import) |
iso |
string | source | ISO URL (CDI import) |
bootc |
string | source | bootc image URI (needs bootc plugin) |
pvc |
string | source | existing PVC name |
sshKey |
string | bootc only | SSH public key baked into the disk |
node |
string | no | schedule on a specific node |
cloudInit |
string | no | cloud-init user-data YAML |
instancetype |
string | no | cluster instancetype name |
preference |
string | no | cluster preference name |
Exactly one source field must be set.
Response (201 Created):
{"name": "web", "namespace": "tailvm"}Bootc response (202 Accepted, build is async):
{"task": "bootc-web-1712345678"}Poll the task ID at GET /api/tasks/{id}.
Get the raw JSON of the VM manifest (the kubevirt VirtualMachine object).
Execute an action on a VM. Valid actions:
| Action | Meaning |
|---|---|
start |
Power on |
stop |
Graceful shutdown |
restart |
Restart |
pause |
Freeze (kubevirt only) |
unpause |
Resume (kubevirt only) |
migrate |
Live-migrate (kubevirt only) β prefer the dedicated endpoint below |
Response: {"status": "ok"}
Live-migrate as a tracked background task with progress in the activity panel. The trigger runs synchronously (so "not migratable / cross-vendor" errors come back immediately); then the server watches the migration to completion.
Body: {"targetNode": "bihar"} (optional β empty lets the scheduler choose,
from among same-CPU-vendor nodes).
Response: 202 {"task": "migrate-β¦"} β poll GET /api/tasks/{id}.
Add or remove a tag, persisted as a corral.dev/tag.<name> VM label (so tags
survive round-trips and are kubectl get vm -l-selectable). GET /api/vms
shows the tags on every VM.
Body: {"tag": "prod", "on": true}
Response: {"tag": "prod", "on": true}
Delete a VM, its PVCs, DataVolumes, hotplug disks, snapshots, proxy resources, and registry entry.
Response: {"status": "deleted"}
Proxmox-style pet pods β plain Kubernetes pods, not KubeVirt VMs. See
docs/adr/0005-containers-as-pet-pods.md for the design.
List all Containers across every namespace.
Response: 200
[{"name": "devbox", "namespace": "corral-vms", "node": "karnataka",
"phase": "Running", "ready": true, "image": "debian:bookworm",
"cpu": 2, "mem": "2Gi", "privileged": true}]phase is the pod phase, or "Stopped" when no pod exists (the CT's
durable identity is its data PVC, which survives Stop). node is empty
when stopped/unscheduled.
Create a Container.
Body:
{"name": "devbox", "namespace": "corral-vms", "image": "docker.io/library/debian:bookworm",
"cpu": 2, "mem": "2Gi", "disk": "10Gi", "storageClass": "", "privileged": true}privileged: true seeds a persistent, mutable root filesystem onto the
data volume and chroots into it (distrobox-style β package installs and
dotfiles survive Stop/Start). false (default) mounts the volume at
/data only.
Response: 201 {"name": "devbox", "namespace": "corral-vms"}
action is start or stop. Stop deletes the pod but keeps the data
volume; Start recreates the pod from the spec recorded on the volume's
annotation.
Response: {"status": "ok"}
Delete a Container: its pod, Service (if it has a tailnet proxy), and data volume.
Response: {"status": "deleted"}
Change CPU count and/or memory. Body: {"cpu": 4, "mem": "8G"}. Live-hotplugs
when the VM is live-migratable; otherwise does a stopβpatchβstart.
Response: {"status": "ok"}
Hotplug a new disk. Body: {"size": "10Gi"}.
Response: {"pvc": "web-disk-2"}
Detach a hotplugged disk. {vol} is the PVC name.
Response: {"status": "removed"}
Expand an existing PVC. Body: {"pvc": "web-disk-1", "size": "40Gi"}. Needs a
StorageClass with allowVolumeExpansion: true.
Response: {"status": "ok"}
Add a secondary NIC. Body: {"nad": "lan", "iface": "eth1"}. Needs Multus + a
NetworkAttachmentDefinition in the VM's namespace.
Response: {"status": "ok"}
List snapshots for a VM. Needs the Snapshot feature gate.
Response: array of {name, phase, creationTime}
Create a snapshot. Body: {"name": "before-upgrade"} (optional; auto-named).
Response: {"name": "before-upgrade"}
Restore a VM to a snapshot. The VM is shut down during restore.
Response: {"status": "restoring"}
Delete a snapshot.
Response: {"status": "deleted"}
Clone a VM (definition and disks). Stop the VM first. Body:
{"target": "web-clone"}. Needs the Snapshot feature gate + a
VolumeSnapshotClass.
Response: {"target": "web-clone"}
Mark or unmark a VM as a template. Body: {"on": true}.
Response: {"isTemplate": true}
The built-in catalog of OS images β curated, ready-to-boot containerdisks.
Response (array):
[
{
"name": "fedora",
"description": "Fedora 42 cloud",
"containerDisk": "quay.io/containerdisks/fedora:42",
"defaultUser": "fedora"
}
]The response then adds the user-defined custom sources, each with the flag "custom": true.
The user-defined custom sources only (for management UI). Same entry shape as
/api/images, all "custom": true. Persisted in the corral-sources
ConfigMap so they survive web-pod restarts.
Add or replace a custom source (idempotent by name).
{"name": "my-image", "kind": "containerDisk", "uri": "ghcr.io/me/img:tag", "description": "optional"}kind is one of containerDisk (boots directly), url (qcow2/raw, CDI
import), or iso (installer ISO). Response: {"status": "ok", "name": "my-image"}.
Remove a custom source. Response: {"status": "removed"}.
List imported images (CDI DataVolumes). These are ISO/qcow2/raw images imported from URLs.
Response: array of {name, namespace, size, phase, progress, source}
Start a CDI import. Body:
{"name": "jammy", "namespace": "tailvm", "url": "https://.../jammy.qcow2", "size": "10Gi"}Response: {"name": "jammy"}
Delete an imported image (DataVolume).
Response: {"status": "deleted"}
Poll a bootc build task. Returns live build log.
Response:
{"status": "running", "log": "Pulling image...\nInstalling...", "error": ""}Statuses: "running" β "done" (success) or "error" (failure).
Guest-agent info (OS, hostname, users, filesystems). Returns 503 if the
agent isn't reachable.
Recent K8s events for the VM (Proxmox-style task log).
Live CPU/memory/disk metrics. Returns null values if the metrics-server isn't available.
Retained usage samples for the Summary-panel charts. The server samples every
live VM every ~15s into a bounded in-memory ring buffer (~1h). Returns
[{"t": <epoch-ms>, "cpu": <millicores>, "mem": <bytes>}, β¦] β an empty array
when metrics-server is absent. The mem field is absent when metrics-server
gives no memory value.
The same samples, added up over all VMs that run now. The Datacenter dashboard charts show them.
The same samples, added up over the VMs that run now on one node. The node dashboard charts show them.
The newest sample of each VM that runs now, busiest first. The "top VMs" dashboard widgets use it.
Query: by=cpu (default) or by=mem sets the sort order. node=<name>
keeps only the VMs on that node. limit=<n> keeps the first n rows.
Response (array):
[{"namespace": "corral-vms", "name": "db-prod", "node": "corral-2", "cpu": 900, "mem": 4294967296}]The caller's tailnet identity and privilege, for the UI to show the logged-in
user and switch to read-only for non-admins. Identity comes from the Tailscale
ingress headers; CORRAL_ADMINS controls admin (see
ADR-0003).
Response: {"login": "alice@github", "name": "Alice", "admin": true, "enforced": false}
When an allowlist exists, the server rejects requests that change state
(non-GET) from non-admins with 403.
List cluster nodes with readiness, roles, kubelet version, and architecture.
Response (array):
[{"name": "bihar", "ready": true, "roles": "control-plane,master", "kubelet": "v1.36.1", "arch": "amd64"}]Cluster capability flags β the UI gates controls on these.
Response:
{
"storageClass": "longhorn",
"canExpand": true,
"canSnapshot": true,
"bootc": false
}| Field | Meaning |
|---|---|
storageClass |
Default StorageClass for new VM disks |
canExpand |
Whether allowVolumeExpansion: true exists |
canSnapshot |
Whether a VolumeSnapshotClass is available |
bootc |
Whether the bootc plugin is compiled in |
Cluster instancetypes and preferences for the create wizard.
Response:
{"instancetypes": ["u1.medium", "u1.large"], "preferences": ["fedora", "ubuntu"]}Multus NetworkAttachmentDefinitions available for secondary NICs.
Run all Corral cluster diagnostics. Returns a list of checks (name, status, message).
Run the fixable diagnostics (e.g. create missing namespaces).
Response: {"fixed": ["created namespace tailvm"]}
Marketplace entries merged with locally-installed state.
Response (array):
[{"name": "bootc", "description": "...", "version": "0.1.0", "installed": false, "inStore": true}]Install a plugin from the marketplace.
Response: {"installed": "bootc"}
Remove an installed plugin.
Response: {"status": "removed"}
Upgrade to WebSocket β bridges virtctl vnc --proxy-only. Use with noVNC in
the browser.
Upgrade to WebSocket β bridges virtctl console. Use with xterm.js in the
browser.
Download a VM disk backup. Stop the VM first (the RWO disk is busy while
the VM runs). Triggers virtctl vmexport and streams the result.
| Query | Format | Content-Type |
|---|---|---|
| (default) | gzipped raw (name.img.gz) |
application/gzip |
?format=qcow2 |
compressed qcow2 (name.qcow2) |
application/octet-stream |
qcow2 needs qemu-img on the server (it converts the raw export); if absent the
request returns 501 and the default raw.gz still works.
Pools are the user-defined folders of ADR-0008. Paths travel in the body or the query string, never in the route, because a nested path contains slashes.
| Method | Route | Notes |
|---|---|---|
GET |
/api/folders |
Tree with members resolved against the live fleet, plus unfoldered |
POST |
/api/folders |
{path} β nesting with /, max depth 8 |
DELETE |
/api/folders?path=β¦ |
Members are unfoldered, never deleted |
POST |
/api/folders/move |
{from, to} β re-parent a pool |
POST |
/api/folders/members |
{path, ref} β what a drag-and-drop lands on |
DELETE |
/api/folders/members?ref=β¦ |
|
POST |
/api/folders/action |
{path, action} β fans out with per-member results |
Moving an instance between backends (ADR-0010) is two endpoints, because the
decision and the doing are separate. A move stops the guest β it is not a
live migration, and corral migrate remains the live, same-backend operation.
| Method | Route | Notes |
|---|---|---|
GET |
/api/move/destinations |
Which backends can receive, and why the others cannot |
POST |
/api/move/preflight |
Returns the plan. Touches nothing β safe to call on every drag |
POST |
/api/move |
Commits, after re-running the preflight server-side |
Body for both: {ref, toBackend, toContext?, toNamespace?, name?, scratch?, deleteSource?}.
A refused preflight is a 200 β the refusals are the answer. An error
status would make a UI show "request failed" where it should show three reasons.
A refused commit is a 409 with the same list. The move stops the source and
keeps it unless deleteSource is set.
| Method | Route | Notes |
|---|---|---|
GET |
/healthz |
Liveness probe endpoint. Returns 200 OK (ok\n) when HTTP server is responding |
GET |
/readyz |
Readiness probe endpoint. Returns 200 OK ({"status":"ready"}) if the registry store is initialized; returns 503 Service Unavailable if unready |
| Method | Route | Notes |
|---|---|---|
GET |
/metrics |
Prometheus text exposition. Requires corral web --metrics |
Served from a cached snapshot refreshed on a background timer, so a scrape never
fans out to the backends. Corral publishes corral_collection_age_seconds for that
reason: without it, a frozen collector is indistinguishable from a stable fleet.
The context label is the configured context's name, so the instance series
and corral_backend_up join. The pool label is the one that spans backends β
sum by (pool) (corral_instance_running) answers "is my application stack up"
across a KubeVirt cluster and a Proxmox host at once. See ADR-0011 for the full
series list and label rationale.
Without --metrics the endpoint still answers 200 with
corral_collection_success 0. A scraper that receives a 503 records nothing,
and "up but no collection" is the state that deserves an alert.
The server serves the embedded SPA at the root:
| Route | File |
|---|---|
/ |
index.html |
/style.css |
style.css |
/app.js |
app.js |
/icons.js |
icons.js |
All other paths fall through to the Go 1.22 http.ServeMux 404.