Inference · Python
LLMLoad
Open SourceLLMLoad is a launcher and load-testing utility for current llama.cpp builds. It makes it easy to run multiple models on mixed GPUs — Intel Arc Pro, ROCm, CUDA — with one command, then load-test token throughput against the running server.
github.com/AStarStarship/… →Specification
| Language | Python + shell |
|---|---|
| Backends | Vulkan (Intel Arc), ROCm, SYCL, CUDA, hybrid |
| Upstream | llama.cpp (fast-forward-only updates) |
Design & implementation
One-command lifecycle
UPDATE, BUILD, and LS are subcommands. Launch uses positional arguments for port, backend, devices, model, and runtime settings. Record the printed revision when comparing runs; upstream llama.cpp changes can still regress.
Mixed-GPU builds
BUILD HYBRID produces a llama.cpp that exposes Arc cards through Vulkan and NVIDIA cards through CUDA in the same binary. Device numbers come from LS — never assume PCI ordering.
Token-throughput load testing
llmload.py hammers the running server with configurable concurrency, request count, warmup, and max-tokens — real throughput numbers, not vibes.
Context-budget testing
Configure the context budget for the model and hardware, then measure memory use and token throughput. Long-context capacity depends on the selected build and available memory.
Related