LogsterClassifier — thin vLLM image + external model (air-gapped)
This is the model-less delivery of the LogsterClassifier. Instead of one ~80–90 GB image with the weights baked in, you receive:
- a thin vLLM image (~11 GB) that you load into your registry like any other Logster image, and
- the model weights as a plain tarball that you upload to an OpenShift PVC.
At run time the PVC is mounted into the vLLM pod at /models, and vLLM serves
the weights from there. This keeps the image small enough to push to Quay
without the large-layer upload problems the baked-in image causes.
The image is built from vLLM with
a static self-signed cert, OpenShift arbitrary-UID hardening, HF_HUB_OFFLINE=1,
and a /models volume. Its CMD loads the model from /models/logster-26b-a4b-it,
so the weights must land at exactly that path inside the PVC.
What's in the delivery
A single .zip containing:
| File | What it is |
|---|---|
vllm-eunomatix:logster-26b-a4b-it-1-0-2.docker.tar.gz |
The thin vLLM image (docker save format). |
images-manifest.txt |
image-ref archive-filename mapping for the image above. |
sha256sum.txt |
Checksums for the files in this zip. |
logster-26b-a4b-it-1-0-2.tar |
The model weights (config.json, tokenizer, *.safetensors, …). Unpacks to a top-level logster-26b-a4b-it/ directory (~52 GB; it also carries a small .cache/ of Hugging Face download metadata — harmless, leave it). |
ubi-9.docker.tar.gz |
A small ubi9/ubi utility image (has a shell and tar) used by the model-loader pod in step 3, shipped so the air-gapped cluster never has to pull from the internet. |
All artifacts are linux/amd64.
Use a Linux (or macOS) bastion. The image archive's filename contains a
:, which Windows filesystems reject — extracting the zip on a Windows transfer host fails.
1. Verify and unpack
On the bastion / transfer host:
unzip logster-classifier-vllm-1-0-2.zip
cd logster-classifier-vllm-1-0-2/
sha256sum -c sha256sum.txt # every line must say OK
Integrity is verified after unzipping (
sha256sum.txtcovers the files inside the zip). Ifunzipitself fails, the download is truncated — re-transfer the zip.
2. Load the thin image into your registry
Exactly like the other Logster images (see the Helm install guide, §0.4). Using
podman (swap in docker/skopeo if that's what the bastion has):
podman login quay.example.com # robot account + token
while read -r ref file; do
target="quay.example.com/logster/${ref##*/}"
podman load -i "$file"
podman tag "$ref" "$target"
podman push "$target"
podman rmi "$ref" "$target"
done < images-manifest.txt
Result: quay.example.com/logster/vllm-eunomatix:logster-26b-a4b-it-1-0-2
(confirm the exact tag against images-manifest.txt).
The archive is an OCI-layout image (with a Docker-compatible manifest).
podman loadis the tested path.docker loadneeds Docker ≥ 25; withskopeo, copy from the decompressed archive using theoci-archive:transport (notdocker-archive:as in the Helm guide's §0.4).
The manifest also lists the shipped ubi9/ubi utility image, so the same loop
pushes quay.example.com/logster/ubi:9 — the image the model-loader pod in
step 3 uses. The cluster is air-gapped and cannot pull it from Red Hat's
registry, which is why it ships in the zip. (Don't substitute ubi-minimal:
it has no tar, which oc cp and the extract step require.)
3. Create the PVC and upload the model
The weights live on a ReadWriteOnce PVC that the vLLM pod mounts read-only.
Size it comfortably above the extracted model size (the .tar unpacks to
~52 GB; 120Gi leaves room for the Option-A staging copy below).
Everything in this guide lives in the logster project (the same one the
Helm guide creates in Part 1) — the PVC, the loader pod, and the classifier
Deployment must all be in it, or the pods won't find the claim. The manifests
below pin namespace: logster explicitly, so they are safe to apply from any
active project:
oc project logster # optional with the explicit namespaces, but keeps oc exec/cp short
oc apply -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: logster-classifier-model
namespace: logster
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 120Gi
# storageClassName: <your-default-or-named-class>
EOF
You cannot copy into a PVC that nothing is mounting, so start a tiny loader pod that mounts it, copy the weights in, then delete the pod:
oc apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata:
name: model-loader
namespace: logster
spec:
restartPolicy: Never
imagePullSecrets:
- name: regcred # the pull secret from the Helm guide, Part 1
containers:
- name: loader
# Pushed by the step-2 loop — the cluster can't reach registry.access.redhat.com.
image: quay.example.com/logster/ubi:9
command: ["sleep", "infinity"]
volumeMounts:
- name: model
mountPath: /models
volumes:
- name: model
persistentVolumeClaim:
claimName: logster-classifier-model
EOF
oc wait --for=condition=Ready pod/model-loader -n logster --timeout=120s
Get the weights onto the PVC so they land at /models/logster-26b-a4b-it — the
path the image's CMD loads from. Two ways, both writing into the loader pod's
mounted PVC:
Option A — oc cp the tar in, extract in-pod (simplest). oc cp copies a
file into a running container (it can't target a PVC directly, which is why the
loader pod exists — and it uses tar internally, present in the UBI image):
oc cp logster-26b-a4b-it-1-0-2.tar logster/model-loader:/models/model.tar
oc exec -n logster model-loader -- tar -xf /models/model.tar -C /models
oc exec -n logster model-loader -- rm /models/model.tar # reclaim the space
Option B — stream-extract without staging the tar on the PVC (avoids needing 2× the space during extraction):
Either way, verify the final layout:
# config.json, tokenizer, *.safetensors must be directly under this dir.
oc exec -n logster model-loader -- ls /models/logster-26b-a4b-it
The commands above assume the tar unpacks to a top-level
logster-26b-a4b-it/directory. If it extracts its files at the top level instead, extract into an explicit target:oc exec -n logster model-loader -- mkdir -p /models/logster-26b-a4b-itthen extract with-C /models/logster-26b-a4b-it. The end state must be/models/logster-26b-a4b-it/config.jsonetc.
Then remove the loader pod — the PVC keeps the data:
4. Run the classifier pod against the PVC
The thin image is a self-contained vLLM server: give it a GPU (≥ 80 GB
VRAM for the single-GPU default), mount the model PVC read-only at
/models, and expose port 8000. No Helm chart is involved — these three
manifests (Deployment, Service, Route) are all it needs.
Apply them, adjusting the image reference, GPU nodeSelector/tolerations, and storage to your cluster:
oc apply -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata:
name: logster-classifier
namespace: logster
labels: {app: logster-classifier}
spec:
replicas: 1
# Only one pod can hold the GPU — recreate rather than roll.
strategy: {type: Recreate}
selector:
matchLabels: {app: logster-classifier}
template:
metadata:
labels: {app: logster-classifier}
spec:
imagePullSecrets:
- name: regcred
# Pin to a GPU node and tolerate the GPU taint if present.
nodeSelector:
nvidia.com/gpu.present: "true"
tolerations:
- {key: nvidia.com/gpu, operator: Exists, effect: NoSchedule}
containers:
- name: classifier
image: quay.example.com/logster/vllm-eunomatix:logster-26b-a4b-it-1-0-2
imagePullPolicy: IfNotPresent
ports:
- {containerPort: 8000, name: https}
resources:
requests: {cpu: "2", memory: 16Gi, nvidia.com/gpu: 1}
limits: {cpu: "8", memory: 32Gi, nvidia.com/gpu: 1}
volumeMounts:
- {name: model, mountPath: /models, readOnly: true}
# vLLM needs a large /dev/shm; back it with a memory emptyDir so
# the pod stays restricted-v2 (no hostIPC).
- {name: shm, mountPath: /dev/shm}
# The server serves HTTPS on 8000 and loads weights before it
# answers — the startupProbe owns a generous boot window.
startupProbe:
httpGet: {path: /health, port: 8000, scheme: HTTPS}
periodSeconds: 15
failureThreshold: 120 # up to ~30 min to load
readinessProbe:
httpGet: {path: /health, port: 8000, scheme: HTTPS}
periodSeconds: 15
volumes:
- name: model
persistentVolumeClaim:
claimName: logster-classifier-model
- name: shm
emptyDir: {medium: Memory, sizeLimit: 16Gi}
---
apiVersion: v1
kind: Service
metadata:
name: logster-classifier
namespace: logster
spec:
selector: {app: logster-classifier}
ports:
- {name: https, port: 8000, targetPort: 8000}
---
# Only needed if something OUTSIDE the cluster must call the endpoint (e.g. the
# inference service runs in another cluster). TLS passthrough preserves the
# server's self-signed cert end to end. Omit for in-cluster-only access.
apiVersion: route.openshift.io/v1
kind: Route
metadata:
name: logster-classifier
namespace: logster
spec:
to: {kind: Service, name: logster-classifier}
port: {targetPort: 8000}
tls: {termination: passthrough}
EOF
Wait for the model to load (first pull is now small since the weights aren't in the image; the VRAM load still takes a few minutes):
oc rollout status deploy/logster-classifier -n logster --timeout=40m
oc get pods -l app=logster-classifier -n logster # READY 1/1 when serving
READY 1/1 means the server passed its /health check — the model is loaded
and answering.
5. Capture the endpoint and install the application chart
The inference service calls the classifier's OpenAI-compatible endpoint. How you address it depends on where the inference service runs:
- Same cluster (typical): use the in-cluster Service —
https://logster-classifier.logster.svc:8000/v1/chat/completions. You can skip the Route entirely. - Separate cluster: use the Route host —
oc get route logster-classifier -n logster \
-o jsonpath='https://{.spec.host}/v1/chat/completions{"\n"}'
Then install the application chart (Helm guide, Part B) pointing
inference.llmEndpoint at that URL and inference.modelName at
eunomatix/logster-26b-a4b-it (the --served-model-name the image advertises).
Troubleshooting
- Pod crashes with "model not found" / HF offline error — the weights aren't
at
/models/logster-26b-a4b-it. Re-mount the PVC in a loader pod and check the path; the image runsHF_HUB_OFFLINE=1, so it never falls back to a download. oc exec ... tarfails on the classifier pod — the classifier image is distroless-style with no shell/tar. Use the loader pod above (a UBI image) for all copy operations, never the classifier pod.- Permission denied writing to
/models— the loader pod must be able to write the PVC. Underrestricted-v2SCC the mount is group-0 writable; if your storage backend enforces different ownership, run the loader with an fsGroup matching the PVC.