Running models locally
Learn how to run models locally with the WriftAI CLI
Overview
wriftai model start runs a model locally, and wriftai model serve runs a model server for local
development. Both currently use Docker to do this.
Requirements
- Docker, version 29.6.0 or later
GPU support
Both commands can additionally expose GPUs to a model. This depends on Docker's NVIDIA GPU integration.
Prerequisites
To run a model locally with GPU acceleration, you also need:
- An NVIDIA GPU
- NVIDIA drivers installed on the host
- The NVIDIA Container Toolkit installed and configured for Docker
If any of these are missing, the model runs on CPU instead of failing outright.
You can confirm your setup is working independently of wriftai by running:
docker run -it --rm --gpus all ubuntu nvidia-smiIf this prints a GPU table, your host is ready. If it errors, the issue is in your Docker/NVIDIA setup rather than in wriftai.
Choosing which GPUs to expose
Both wriftai model start and wriftai model serve accept a --gpus flag. All available GPUs
are exposed by default.
For the full list of accepted --gpus values and usage examples, see the
model start and model serve
reference pages.
Caching
Both commands mount a host directory into the container to persist caches across restarts. This
covers PyTorch's torch.compile cache and Hugging Face model downloads. Without this mount, a
model that compiles with torch.compile would recompile on every restart, and a model that
downloads weights at runtime would re-download them each time.
Caches are stored under ~/.cache/wriftai, with a separate subdirectory per model, so one
model's cache can never be corrupted by or interfere with another's. Each model's subdirectory is
mounted into its container as /tmp (model containers otherwise run with a read-only
filesystem). The following environment variables are set by default to direct caches there:
| Variable | Default | Purpose |
|---|---|---|
TORCHINDUCTOR_CACHE_DIR | /tmp/torchinductor | torch.compile graph and autotune cache |
TRITON_CACHE_DIR | /tmp/triton | Compiled Triton kernel cache |
HF_HOME | /tmp/huggingface | Hugging Face model/tokenizer downloads |
These can be overridden with --env, for example to point at a different location or disable a cache:
wriftai model start owner/name --env TORCHINDUCTOR_CACHE_DIR=/tmp/custom-pathCache growth is currently unbounded. Periodically clearing ~/.cache/wriftai is safe if disk
usage becomes a concern; to clear a single model's cache instead of everything, remove its
subdirectory.