E.g. Debian with Curl, Ubuntu Build etc.
  • Python 74.2%
  • HTML 21.7%
  • Shell 4.1%
Find a file
john enlow 406727496b
Some checks failed
Build-Publish-GPU / build-gpu (push) Successful in 6s
Build-Publish-Multi-Arch / build (linux/arm64) (push) Successful in 13s
Build-Publish-Multi-Arch / build (linux/amd64) (push) Successful in 16s
Build-Publish-Multi-Arch / create-manifest (push) Successful in 8s
Build-Publish-GPUBig / build-gpubig (push) Has been cancelled
ci: add weekly schedule for Docker image rebuilds
2026-09-21 11:19:43 +12:00
.gitea/workflows ci: add weekly schedule for Docker image rebuilds 2026-09-21 11:19:43 +12:00
attic Retire 15 unused base images to attic/; cap CUDA bumps at sm<7.5 2026-08-24 19:56:09 +12:00
.gitignore Add base-images.py to check and refresh upstream base images 2026-08-24 19:17:42 +12:00
bake_pocket.py Move pocket-tts bake script to a file; heredoc + --mount crashed build 2026-07-04 23:08:31 +12:00
base-images.py Fix CI 429s: authenticate Docker Hub pulls, cache registry reads 2026-08-24 20:03:50 +12:00
Dockerfile.accelerated_base Bump base images: CUDA 12.9.2, Python 3.14, refresh accelerated_base 2026-08-24 21:58:34 +12:00
Dockerfile.clabtree-api-base clabtree-api-base: bump lxml to 6.1.2 for Python 3.14 2026-08-24 23:03:03 +12:00
Dockerfile.clabtree-imagetools-base Fix imagetools-base build: guard the APT proxy on the host resolving 2026-08-24 20:21:41 +12:00
Dockerfile.clabtree-music-pascal Pin diffusers <0.38.0 to fix infer_schema crash on PyTorch 2.5.1 2026-05-05 22:57:57 +12:00
Dockerfile.clabtree-tool-base Bump base images: CUDA 12.9.2, Python 3.14, refresh accelerated_base 2026-08-24 21:58:34 +12:00
Dockerfile.clabtree-youtube-tool-base ci: add weekly schedule for Docker image rebuilds 2026-09-21 11:19:43 +12:00
Dockerfile.knowledgebase-base knowledgebase-base: remove stale apt proxy file before apt-get 2026-05-10 23:35:18 +12:00
Dockerfile.pytorch-gpu pytorch-gpu: drop Flash Attention 2026-08-24 23:43:58 +12:00
Dockerfile.pytorch-gpu-pascal Bump base images: CUDA 12.9.2, Python 3.14, refresh accelerated_base 2026-08-24 21:58:34 +12:00
package-lock.json add package-lock.json 2026-05-25 15:49:39 +12:00
README.md Update README to note redundant rm -f proxy lines in base images 2026-08-30 13:11:36 +12:00

Generic Docker Images

Reusable Docker base images published to forge.jde.nz/public/.

Requires forgejo runners gpu, gpubig to run the CI.

Available Images

pytorch-gpu

PyTorch + CUDA base image for NVIDIA GPU workloads. Supports all NVIDIA GPUs from GTX 1060 to RTX 5090.

GPU Compute Capabilities: 6.1 (GTX 10xx), 7.5 (RTX 20xx), 8.0 (RTX A6000/A40), 8.6 (RTX 30xx), 8.9 (RTX 40xx), 9.0 (RTX PRO 6000), 10.0 (RTX 50xx/Blackwell)

Includes:

  • CUDA 12.8.1 runtime
  • PyTorch 2.7+ with torchaudio
  • transformers, accelerate, safetensors
  • FFmpeg, libsndfile, sox (audio processing)
  • Python 3 venv at /opt/venv

Usage:

FROM forge.jde.nz/public/pytorch-gpu:latest

RUN pip install my-project-specific-packages
COPY my_app/ /app/
CMD ["python", "/app/server.py"]

Architecture: x86_64 only (CUDA)

Size: ~8-10GB

No Flash Attention. Removed 2026-08-24 — no consumer imported it, and FA2 requires Ampere (sm_80) or newer so it could not serve the 6.1/7.5 architectures this image targets. It was also the slowest and most brittle step in the repo. See the comment in Dockerfile.pytorch-gpu before re-adding it. Consumers needing fused attention should use PyTorch SDPA, or install what they need themselves (as clabtree-upscale does with SageAttention).

clabtree-imagetools-base

Pre-built base image for clabtree-imagetools. Layers all heavy ML deps on top of pytorch-gpu.

Includes: torchvision, transformers 4.x (pinned below 5.0 for Flux2KleinPipeline compat), diffusers from git (Flux2KleinPipeline), torchao, onnxruntime-gpu, insightface, gfpgan, realesrgan, opencv-python-headless, pillow-heif

Purpose: Avoids 3+ minute pip builds on every imagetools deploy. Consumer image just copies app code on top.

Architecture: x86_64 only (CUDA, GPU runner only)

clabtree-api-base

Pre-built base image for clabtree-api. Bundles all heavy Python dependencies so API deploys only need to copy app code.

Includes:

  • PyTorch CPU (no CUDA — API runs on CPU-only server)
  • pipecat-ai with voice pipeline extras (silero, webrtc, openai, smart-turn)
  • LiveKit agents + plugins, sherpa-onnx, LocalVQE neural AEC (compiled from source)
  • fastembed (ONNX embeddings)
  • weasyprint + pandoc (PDF generation)
  • ffmpeg, audio/graphics system libraries
  • All other API Python deps (fastapi, httpx, aiosqlite, etc.)
  • All models pre-baked so the API image and runtime download nothing: livekit turn-detector EOU models (via download-files), sherpa STT/TTS, LocalVQE GGUF, and the BAAI/bge-small-en-v1.5 fastembed model. Cache dirs are pinned (HF_HOME, FASTEMBED_CACHE_PATH) so baked models resolve at runtime.

Purpose: Avoids per-deploy HuggingFace fetches. The consumer image (clabtree-api) is pure app code: FROM base + COPY . ..

Architecture: x86_64 only

accelerated_base

Media processing base image with Intel QuickSync / VA-API + FFmpeg hardware acceleration.

Features:

  • Debian trixie (Debian 13) — ships the Intel iHD media driver 25.x, which supports modern Intel iGPUs including Arrow Lake-S. (Ubuntu 22.04's iHD 22.x did not, so VA-API hardware encode failed on Arrow Lake.)
  • FFmpeg (Debian 7.1, VA-API enabled) + FFmpeg dev libraries
  • Intel QuickSync / VA-API via intel-media-va-driver-non-free (x86_64)
  • Python venv at /opt/venv (on PATH) — consumer images pip install directly, no PEP 668 friction
  • OpenCV (headless), numpy, Pillow, imageio; non-root appuser

Usage: layer your app on top and pass --device /dev/dri at runtime for hardware encode/decode:

FROM forge.jde.nz/public/accelerated_base:latest
USER root
RUN pip install --no-cache-dir my-deps
COPY app/ /app/

Architecture: x86_64 (Intel QuickSync) + arm64 (VA-API libraries only — no Intel iGPU driver, software fallback)

Note: the previous Ubuntu CUDA-11.8 NVENC libs (libnvidia-encode-515 etc.) were removed — they were unused by consumers. Use a CUDA base image if you need NVIDIA NVENC.

clabtree-tool-base

Shared floor for the small FastAPI tool services. python:3.13-slim plus the only four packages all of them need:

  • fastapi
  • starlette>=1.0.1
  • uvicorn[standard]
  • httpx

Installed once so the 17 tools share one 54MB layer instead of each building its own — pip installs are not reproducible, so identical requirements.txt were still producing different layers.

Usage:

FROM forge.jde.nz/public/clabtree-tool-base:latest

WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt   # only the extras
COPY . .

Leave the four core packages out of the template's requirements.txt — they are already here. Anything else (python-multipart, pydantic, pillow, …) stays in the template.

Retired images (attic/)

These had no remaining consumer anywhere in the tree, so as of 2026-08-24 they live in attic/ and CI no longer builds them. The Dockerfiles are kept for reference; the images already published to forge.jde.nz/public/ were left in place, so anything still pulling one keeps working — it just stops being refreshed.

Between them they accounted for most of the repo's build time: the nemotron-speech-blackwell chain alone is three images of PyTorch, vLLM, triton and mamba-ssm compiled from source for Blackwell (README used to quote "2-3 hour build"), and orthrus-base compiles Flash Attention 2 for five architectures.

Image Why
nemotron-speech-blackwell, -torch, -vllm built for clabtree-parakeet-asr, which no longer exists — streaming STT is clabtree-parakeet-stt on its own image
nemotron-speech-ada same consumer
qwen3-asr-base no consumer template was ever kept
orthrus-base clabtree-llm-orthrus destroyed 2026-05-16
clabtree-faceswap-intel-base, -nvidia-base HyperSwap templates deleted; face swap is BFS-LoRA inside clabtree-imagetools
ik_llama.cpp-server-cuda, llama-cpp-server-sycl every llm-* worker uses ghcr.io/ggml-org/llama.cpp:server-cuda directly
comfyui-cuda only consumer was clabtree-video (Wan 2.2), retired
clabtree-minimax-music3-base clabtree-music-5090-minimax-music3 never deployed
debian-curl, ds-network-sidecar no references anywhere, including dropshell itself
llm-cpu-gemma-chat standalone demo image; baked a 14GB GGUF at build time

To bring one back: git mv attic/Dockerfile.<name> . and restore its build step in the matching workflow.

Registry

Images are available at:

  • forge.jde.nz/public/<image_name>:latest — multi-arch manifest (or x86_64 for GPU images)
  • forge.jde.nz/public/<image_name>:latest-x86_64 — x86_64 specific
  • forge.jde.nz/public/<image_name>:latest-aarch64 — arm64 specific (non-GPU only)

Building

Images are automatically built and published by CI on push to main (.gitea/workflows/buildpublish*.yaml, split by runner: generic multi-arch, gpu, gpubig). CI only rebuilds an image whose Dockerfile changed — that is what needs_build in each workflow tests — so a base that moved underneath a floating tag will not trigger anything on its own.

One consequence worth knowing: a failed build is never retried automatically. needs_build looks at the last commit's diff, so once the commit that changed a Dockerfile has landed, a later push that doesn't touch that file will skip it — leaving the published image on its old content, and (on the multi-arch job) the two architectures out of step with each other. Force the retry with ./base-images.py --update, which stamps a # base-refresh: line so the file lands in the next diff.

To build one locally: docker build -f Dockerfile.<name> . (the GPU images need a CUDA toolchain and are only practical on the GPU runners).

The APT cache is gone — builds fetch direct

pb-lxcaptcache no longer exists. CI used to pass --build-arg APT_PROXY=http://pb-lxcaptcache:3142 to every build; that has been removed, so apt goes straight to the distro mirrors.

This mattered more than it sounds, because an Acquire::http::Proxy pointing at an unresolvable host does not degrade gracefully. Every index fetch fails, the package lists come back empty, and the build dies on a thoroughly misleading E: Unable to locate package <something that obviously exists>. That is what broke clabtree-imagetools-base — it had not been rebuilt since 2026-04-06, so nothing had exercised it since the cache disappeared.

With no --build-arg, ARG APT_PROXY="" makes the proxy blocks in pytorch-gpu, accelerated_base and clabtree-imagetools-base inert, so none of them needed changing. They are harmless dead code; delete them opportunis- tically next time one of those images has to rebuild anyway, rather than paying for a rebuild just to tidy up.

knowledgebase-base and clabtree-imagetools-base both start with rm -f /etc/apt/apt.conf.d/01proxy. That guarded against the published pytorch-gpu carrying a stale proxy file baked in. It no longer does — pytorch-gpu:latest was rebuilt 2026-08-24 with an empty APT_PROXY, so the file is never written. Those rm -f lines are now redundant rather than protective, and cost nothing to keep. If you ever reintroduce a cache, guard it on the host resolving — accelerated_base has the pattern.

Docker Hub rate limits

Anonymous Docker Hub pulls are capped at 100 manifest reads per hour per source IP — and every runner shares one egress address with each other and with the dev machines. That budget is easy to exhaust by accident, and when it goes, builds fail on the FROM line itself:

ERROR: failed to solve: unexpected status from HEAD request to
https://registry-1.docker.io/v2/library/python/manifests/3.13-slim: 429 Too Many Requests

That is not a broken Dockerfile — it is the shared quota. Check what's left with:

TOKEN=$(curl -s "https://auth.docker.io/token?service=registry.docker.io&scope=repository:ratelimitpreview/test:pull" \
  | python3 -c 'import json,sys; print(json.load(sys.stdin)["token"])')
curl -s --head -H "Authorization: Bearer $TOKEN" \
  https://registry-1.docker.io/v2/ratelimitpreview/test/manifests/latest | grep -i ratelimit

DOCKERHUB_USERNAME and DOCKERHUB_TOKEN are required. Every workflow has a Login to Docker Hub step that fails the job if either secret is missing, or if the credential is rejected. They live at the org level next to DOCKER_PUSH_TOKEN:

https://forge.jde.nz/org/public/settings/actions/secrets

DOCKERHUB_USERNAME is the Docker Hub username, not the email — docker login rejects an email address when the password is a personal access token. The token needs only Public Repo Read; CI pulls public images and never pushes there.

This is deliberately fail-fast rather than best-effort. Falling back to anonymous pulls just moves the failure later and makes it much harder to read: you get either a mid-build 429, or apt reporting it "cannot locate" a package that plainly exists. A missing secret should stop the build at the first step with a message that says so.

Keeping base images current

./base-images.py reports how far behind each image is, and can bump it.

./base-images.py                     # check everything (read-only)
./base-images.py --only pytorch-gpu  # just one or a few
./base-images.py --json              # machine-readable
./base-images.py --update            # interactive: pick what to bump/rebuild
./base-images.py --update --commit   # ...and commit the result

It reports two independent kinds of drift, because neither one catches the other:

  • STALE — the published image no longer sits on its base's current layers. Every layer a FROM contributes is inherited verbatim, so the base's diff_ids must be a prefix of ours; when they stop being one, the base has been rebuilt since we last built. This is the only way to see a floating tag (debian:latest, llama.cpp:server, our own :latest bases) move while the Dockerfile stays byte-identical.
  • OUTDATED — a newer upstream tag exists for a pinned base. No digest check can tell you nvidia/cuda:12.8.1-devel-ubuntu22.04 is several versions behind, because that tag is perfectly fresh.

Only the leading version component is compared, and only against tags with the same flavour and the same precision. So 12.8.1-devel-ubuntu22.04 is offered 12.9.2-devel-ubuntu22.04 but never a silent jump to ubuntu24.04, and 3.12-slim is not "upgraded" to 3.12.8-slim. Newer-tag suggestions need a listable tag API, which in practice means Docker Hub; bases elsewhere (ghcr.io, forge.jde.nz) get the layer check only.

CUDA suggestions are capped by the arch list. CUDA 13.0 removed offline compilation for everything below compute capability 7.5, so a Dockerfile whose TORCH_CUDA_ARCH_LIST still contains 6.1 (both pytorch-gpu and pytorch-gpu-pascal do — that is GTX 10xx support) is only ever offered 12.x. Suggesting 13.x there would produce a build that fails on compute_61.

Only Dockerfile.* in the repo root is checked; attic/ is ignored.

Images are listed parents-first, so bumping pytorch-gpu before its dependants is the obvious order — rebuilding a parent makes every child STALE.

For a STALE image there is nothing in the Dockerfile to change, so --update writes a # base-refresh: <timestamp> comment at the end of the file. That is purely a lever for CI's change detection; it has no effect on the built image.

Requires docker buildx (it reads registries using the credentials already in ~/.docker/config.json) and outbound HTTPS. No pip installs.