Light Logo
CLI (v0.57.3)

Running models locally

Learn how to run models locally with the WriftAI CLI

Overview

wriftai model start runs a model locally, and wriftai model serve runs a model server for local development. Both currently use Docker to do this.

Requirements

  • Docker, version 29.6.0 or later

GPU support

Both commands can additionally expose GPUs to a model. This depends on Docker's NVIDIA GPU integration.

Prerequisites

To run a model locally with GPU acceleration, you also need:

If any of these are missing, the model runs on CPU instead of failing outright.

You can confirm your setup is working independently of wriftai by running:

docker run -it --rm --gpus all ubuntu nvidia-smi

If this prints a GPU table, your host is ready. If it errors, the issue is in your Docker/NVIDIA setup rather than in wriftai.

Choosing which GPUs to expose

Both wriftai model start and wriftai model serve accept a --gpus flag. All available GPUs are exposed by default.

For the full list of accepted --gpus values and usage examples, see the model start and model serve reference pages.

Caching

Both commands mount a host directory into the container to persist caches across restarts. This covers PyTorch's torch.compile cache and Hugging Face model downloads. Without this mount, a model that compiles with torch.compile would recompile on every restart, and a model that downloads weights at runtime would re-download them each time.

Caches are stored under ~/.cache/wriftai, with a separate subdirectory per model, so one model's cache can never be corrupted by or interfere with another's. Each model's subdirectory is mounted into its container as /tmp (model containers otherwise run with a read-only filesystem). The following environment variables are set by default to direct caches there:

VariableDefaultPurpose
TORCHINDUCTOR_CACHE_DIR/tmp/torchinductortorch.compile graph and autotune cache
TRITON_CACHE_DIR/tmp/tritonCompiled Triton kernel cache
HF_HOME/tmp/huggingfaceHugging Face model/tokenizer downloads

These can be overridden with --env, for example to point at a different location or disable a cache:

wriftai model start owner/name --env TORCHINDUCTOR_CACHE_DIR=/tmp/custom-path

Cache growth is currently unbounded. Periodically clearing ~/.cache/wriftai is safe if disk usage becomes a concern; to clear a single model's cache instead of everything, remove its subdirectory.