{"guide":"https://nationalcompute.com/FIRST-JOB.md","base_url":"https://access.nationalcompute.com/first-job/","vendor_rule":"Where a recipe has vendor_variants, apply the file for the cluster's gpu_vendor (GET /api/k8s/cluster); a vendor not listed means apply the default, and a recipe with no default is offered only for the vendors listed.","node_rule":"A recipe of cluster_kind public_research is a shell script for one whole node reached over ssh (a Public Research node; no cluster): fetch it on the node and run it there. It reads the node's GPUs itself, so there is no vendor variant to pick.","recipes":[{"name":"prepare","cluster_kind":"slurm","category":"train","default":"prepare.sh","vendor_variants":{},"description":"Fetches the training project into ~/first-job on the login node, stages the PyTorch container image for the cluster's GPUs and tokenizes the dataset on a compute node. Run once before nanogpt.sbatch."},{"name":"nanogpt-sbatch","cluster_kind":"slurm","category":"train","default":"nanogpt.sbatch","vendor_variants":{},"description":"Trains a GPT-2 sized model on Shakespeare with PyTorch DDP across every GPU of one node, as a Slurm batch job. Needs prepare.sh first."},{"name":"nanogpt-node","cluster_kind":"public_research","category":"train","default":"nanogpt-node.sh","vendor_variants":{},"description":"Trains a GPT-2 sized model on Shakespeare with PyTorch DDP on every GPU of one whole node reached over ssh (a Public Research node), in the stock PyTorch container. Run it on the node: it detaches from the ssh session and logs to ~/first-job/nanogpt.out there."},{"name":"nanogpt","cluster_kind":"k8s","category":"train","default":"nanogpt-job.yaml","vendor_variants":{"amd":"nanogpt-job-amd.yaml"},"description":"Trains a GPT-2 sized model on Shakespeare with PyTorch DDP on every GPU of one node, in the stock PyTorch container. Nothing to build."},{"name":"nanogpt-multinode","cluster_kind":"k8s","category":"train","default":"nanogpt-multinode.yaml","vendor_variants":{"amd":"nanogpt-multinode-amd.yaml"},"description":"The same training as one gang-scheduled Job across two whole nodes."},{"name":"shared-logs-reader","cluster_kind":"k8s","category":"train","default":"shared-logs-reader.yaml","vendor_variants":{},"description":"Tails training logs teed to the shared volume, from the CPU worker — readable while GPU nodes come and go. Needs a cluster with a shared volume."},{"name":"kimi-k3-weights","cluster_kind":"k8s","category":"serve","default":"kimi-k3-weights-to-shared.yaml","vendor_variants":{},"description":"Stages the serving recipe's weights (about 1.5 TB) onto the shared volume from the CPU worker, while the GPU node is still arriving."},{"name":"kimi-k3-serve","cluster_kind":"k8s","category":"serve","default":null,"vendor_variants":{"amd":"kimi-k3-serve-amd.yaml"},"description":"Serves a large open model behind a router: one engine replica on every GPU of one whole node, plus the router on the CPU worker."},{"name":"kimi-k3-pd","cluster_kind":"k8s","category":"serve","default":null,"vendor_variants":{"amd":"kimi-k3-pd-amd.yaml"},"description":"The serving recipe across six whole nodes as prefill/decode Deployments. Run one serving recipe at a time."},{"name":"kimi-k3-loadtest","cluster_kind":"k8s","category":"serve","default":"kimi-k3-loadtest.yaml","vendor_variants":{},"description":"Open-loop load test of the serving endpoint from inside the cluster: sweeps request rates and reports tokens per second, time to first token, and completion rate. About 17 minutes."},{"name":"kimi-k3-loadtest-results","cluster_kind":"k8s","category":"serve","default":"kimi-k3-loadtest-results.yaml","vendor_variants":{},"description":"Reads a past load test's results from the shared volume after the test Job is gone."},{"name":"inkling-weights","cluster_kind":"k8s","category":"serve","default":"inkling-weights-to-shared.yaml","vendor_variants":{"nvidia":"inkling-weights-to-shared-nvidia.yaml"},"description":"Stages the Inkling serving recipe's weights (about 530 GiB) onto the shared volume from the CPU worker, while the GPU node is still arriving."},{"name":"inkling-serve","cluster_kind":"k8s","category":"serve","default":null,"vendor_variants":{"amd":"inkling-serve-amd.yaml","nvidia":"inkling-serve-nvidia.yaml"},"description":"Serves Inkling, Thinking Machines' open-weights model (975B parameters, 41B active, 1M context, reasoning and tool calling), as an OpenAI-compatible endpoint: one engine on every GPU of one whole node, no router."},{"name":"finance-analyst-tune","cluster_kind":"k8s","category":"post-train","default":"finance-analyst-tune.yaml","vendor_variants":{"amd":"finance-analyst-tune-amd.yaml"},"description":"Fine-tunes an open chat model into a financial risk analyst on a folder of synthetic filings (LoRA, every GPU of one node) and writes the tuned model to the shared volume. Needs a cluster with a shared volume."},{"name":"finance-analyst-serve","cluster_kind":"k8s","category":"post-train","default":"finance-analyst-serve.yaml","vendor_variants":{"amd":"finance-analyst-serve-amd.yaml"},"description":"Serves the tuned financial analyst from the shared volume as an OpenAI-compatible endpoint on one whole node. Waits from the CPU worker (holding no GPU node) until finance-analyst-tune has finished, then takes the node the tune freed."}]}