Serve DiffusionGemma-Jev (djev) on Cloud Run
with an NVIDIA RTX PRO 6000 Blackwell GPU, built on
mmastrac/djev.
Built-in demo apps, inspired by:
/snake: mizorewww/laya-coreml/dino: virajbhartiya/laya-vs-jev/tetris: trungdq88/jev-tetris
Follows Cloud Run GPU best practices.
export BUCKET="your-gcs-bucket"
export REGION="us-central1" # Supported RTX PRO 6000 regions: us-central1, europe-west4, asia-southeast1, asia-south2
gcloud storage buckets create "gs://${BUCKET}" --location="${REGION}"
hf download nvidia/diffusiongemma-26B-A4B-it-NVFP4 --local-dir /tmp/dgemma
gcloud storage cp -r /tmp/dgemma/* "gs://${BUCKET}/dgemma/"You have two options for deployment:
Deploy the pre-built image with /dev/shm staging and all vllm serve flags baked in.
gcloud beta run deploy djev-dgemma \
--region="us-central1" \
--image=ghcr.io/taeold/djev-run:latest \
--gpu=1 --gpu-type=nvidia-rtx-pro-6000 --no-gpu-zonal-redundancy \
--cpu=20 --memory=80Gi --no-cpu-throttling \
--concurrency=32 --min-instances=0 --max-instances=1 \
--port=8080 \
--network=default --subnet=default --vpc-egress=all-traffic \
--add-volume=name=weights,type=cloud-storage,bucket="${BUCKET}",readonly=false,mount-options=enable-buffered-read=true \
--add-volume-mount=volume=weights,mount-path=/mnt/gcs \
--startup-probe=httpGet.path=/health,httpGet.port=8080,initialDelaySeconds=5,periodSeconds=2,timeoutSeconds=2,failureThreshold=120Directly deploy the upstream Docker container running pure OpenAI completions.
gcloud beta run deploy djev-dgemma \
--region="us-central1" \
--image=docker.io/vllm/vllm-openai:nightly \
--gpu=1 --gpu-type=nvidia-rtx-pro-6000 --no-gpu-zonal-redundancy \
--cpu=20 --memory=80Gi --no-cpu-throttling \
--concurrency=32 --min-instances=0 --max-instances=1 \
--port=8000 \
--network=default --subnet=default --vpc-egress=all-traffic \
--add-volume=name=weights,type=cloud-storage,bucket="${BUCKET}",readonly=false,mount-options=enable-buffered-read=true \
--add-volume-mount=volume=weights,mount-path=/mnt/gcs \
--startup-probe=httpGet.path=/health,httpGet.port=8000,initialDelaySeconds=5,periodSeconds=2,timeoutSeconds=2,failureThreshold=120 \
--set-env-vars="VLLM_FLASHINFER_MOE_BACKEND=masked_gemm,VLLM_ENABLE_V1_MULTIPROCESSING=0" \
--command="/bin/bash" \
--args="-c","cp -r /mnt/gcs/dgemma /dev/shm/dgemma && exec vllm serve /dev/shm/dgemma --served-model-name djev-dgemma --allowed-origins '[\"*\"]' --trust-remote-code --enforce-eager --language-model-only --attention-backend TRITON_ATTN --kv-cache-memory 2G --max-num-seqs 32 --max-model-len 4096 --diffusion-config '{\"canvas_length\":128}' --override-generation-config '{\"max_new_tokens\":null}'"
# Cloud Run & Storage flags:
# --image=docker.io/vllm/vllm-openai:nightly: serves standard vLLM diffusion natively without a custom Dockerfile
# --no-gpu-zonal-redundancy: required for standard regional RTX PRO 6000 quota
# --no-cpu-throttling: keeps all 20 vCPUs active during weight loading and vLLM scheduling
# --network=default --subnet=default --vpc-egress=all-traffic: streams weights from GCS over Google internal networking (~1.05 GiB/s)
# mount-options=enable-buffered-read=true & cp -r to /dev/shm: prefetches 18 GB safetensors shards sequentially from GCS into RAM before vLLM starts
#
# vLLM cold-start & runtime flags (--set-env-vars / --args):
# VLLM_ENABLE_V1_MULTIPROCESSING=0: runs EngineCore in-process so Python/CUDA modules are not imported twice
# VLLM_FLASHINFER_MOE_BACKEND=masked_gemm: selects the low-latency FlashInfer MoE kernel on Blackwell SM120
# --enforce-eager: skips torch.compile and CUDA graph capture on startup
# --language-model-only: skips loading and profiling the unused SigLIP vision encoder
# --kv-cache-memory=2G: pre-allocates a fixed 2 GiB KV cache, skipping the startup memory-profiling forward pass
# --attention-backend=TRITON_ATTN: uses Triton bidirectional attention required by DiffusionGemma
# --diffusion-config='{"canvas_length":128}': configures the 128-token parallel diffusion canvas# 1. Tokenize query
curl -s https://<your-cloud-run-url>/tokenize \
-H "Content-Type: application/json" \
-d '{"prompt": "Classify: Payment issues\nlabel:"}'{"count":8,"max_model_len":4096,"tokens":[4335,1891,236787,35032,4342,107,2491,236787],"token_strs":null}# 2. Diffusion Read
curl -s https://<your-cloud-run-url>/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "djev-dgemma",
"messages": [
{"role": "user", "content": "Classify: Payment issues\nlabel:"}
],
"max_tokens": 8,
"vllm_xargs": {
"diffusion_seed_canvas": [1, 2, 3, 4],
"diffusion_pinned": [0, 1, 2],
"diffusion_max_steps": 1,
"diffusion_read_only": true
}
}'{"id":"chatcmpl-8929997c074788e6","object":"chat.completion","created":1790206319,"model":"djev-dgemma","choices":[{"index":0,"message":{"role":"assistant","content":" Billing"},"finish_reason":"length"}]}ghcr.io/taeold/djev-run:latest also includes support for the /v1/systemone endpoint:
curl -s https://<your-cloud-run-url>/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": "Customer cannot complete checkout; card declined twice.",
"questions": {
"category": {
"type": "choice",
"instructions": "Classify the support ticket",
"criteria": ["billing", "technical", "account"]
}
}
}'{
"model": "djev-dgemma",
"answers": {
"category": {
"choice": "billing",
"probabilities": {
"billing": 0.942,
"technical": 0.041,
"account": 0.017
},
"confidence": 0.942
}
},
"diagnostics": {
"timing": {
"total_ms": 116.3
}
}
}- Cold start from zero instances: ~47.5s (down from 4m 05s baseline)
- Median response time: 61 ms on server, 117 ms end-to-end (
steps=1, samples=1; 121 ms end-to-end withsamples="auto") - Batch throughput: ~100-123 requests/sec at
concurrency=32
JevBench v1.3.0 (N = 231 public suite):
| Configuration | Overall Accuracy | Composite Score | Median Response Time |
|---|---|---|---|
djev-run (steps=1, samples=1) |
81.4% | 73.4 | 117 ms |
djev-run (steps=1, samples="auto") |
81.8% | 73.5 | 121 ms |
api.djev.dev |
81.8% | 73.0 | 237 ms |
1 NVIDIA RTX PRO 6000 GPU (20 vCPU, 80 GiB RAM) costs $3.19 per hour while
active and scales to $0 when idle with --min-instances=0.