Two engines, the same models on the same GPU
basal-rs runs Basal models. We compare its performance with the basal-serve v1.5.0, the Basal model author's reference server in Python and PyTorch. NVIDIA fast mode uses BF16, torch.compile and CUDA graphs. On Macs it uses MLX., the model author's server, on the same GPU. Both engines use 16-bit precision: basal-rs uses f16 and upstream uses BF16. We compare both engines' decisions with full-precision FP32 calculations, the precision used by the author to validate the models. These results describe the engine; decision accuracy comes from the Basal models.
- 900/900 decisions matching the full-precision FP32 reference (basal-1.5-max, A100, f16)
- 0 differences in responses to the same request under load
- 1.5–2.8× more requests on the same GPU (14 of 15 GPU/model pairs)
- 1.3–2.8× less energy per decision across all 15 GPU/model pairs
How we measure
Performance and compatibility results come from reports in the basal-rs repository, with data and test conditions. Both engines are always measured on the same GPU, with basal-rs using its default installation settings.
basal-rs, f16
f16 weights and intermediate results, f32 accumulation in matrix multiplication
upstream, BF16
PyTorch with torch.compile and CUDA graphs; MLX in bf16 on Macs
upstream, FP32
The full precision used by the author to evaluate the models. Used to compare decisions, not performance. For FP32 computation in basal-rs, set dtype: f32 to get 900 / 900.
Single decision
Time for one decision when the server has no other work: 44 sample questions, each in both option orders. Median of 39 measurements after discarding 5.
HTTP server
A running server with 1, 8 or 32 concurrent clients; each request contains one question. The same program sends requests to both servers, and we read energy usage from the GPU.
Mixed load
Documents from 100 to 16k tokens and 1 to 14 questions per request. Only basal-rs was measured: upstream handles about 17 such requests per minute on the max model, so measuring it would take over an extra hour per GPU.
Compatibility
We compare decisions with full-precision FP32 calculations, the precision used by the model author for validation. Both engines must construct exactly the same model input, token for token.
| GPU | Architecture, Compute capability: the NVIDIA architecture version that determines which instructions a GPU supports. | Memory | Power limit | Number of virtual CPU cores (vCPUs) on the host machine. |
|---|---|---|---|---|
| H100 80 GB | Hopper, 9.0 | 80 GB | 700 W | 16 vCPU |
| A100 80 GB | Ampere, 8.0 | 80 GB | 400 W | 16 vCPU |
| L40S | Ada, 8.9 | 48 GB | 350 W | 12 vCPU |
| RTX 6000 Ada | Ada, 8.9 | 48 GB | 300 W | 12 vCPU |
| RTX A6000 | Ampere, 8.6 | 48 GB | 300 W | 6 vCPU |
Five GPUs, three models
The controls above the chart change the model and measurement. The number marked × shows how much better the winning engine's result is for each GPU. When upstream wins, its name precedes the number.
Chart data table
| Measurement | upstream v1.5.0, BF16 (fast) | basal-rs, f16 | Unit |
|---|---|---|---|
| H100 80 GB | 56 | 85 | requests/s (higher is better) |
| A100 80 GB | 24 | 41 | requests/s (higher is better) |
| L40S | 18 | 36 | requests/s (higher is better) |
| RTX 6000 Ada | 10 | 27 | requests/s (higher is better) |
| RTX A6000 | 10 | 21 | requests/s (higher is better) |
In 14 of 15 GPU/model pairs, basal-rs handles 1.5–2.8 times as many requests as upstream. The exception is basal-1.5-mini on H100, where the two servers are similar (basal-rs reaches 0.95 times upstream's result).
Without Python or PyTorch
The chart shows the disk space required to install everything needed to serve Basal models on an NVIDIA GPU. Model files and the GPU driver are excluded. In every mode, upstream runs as basal-serve in Python with torch and transformers, including when Ollama or llama.cpp performs the computation.
Chart data table
| Measurement | upstream: basal-serve + engine | basal-rs 0.1.6 | Unit |
|---|---|---|---|
| PyTorch (fast mode) | 7.08 | 1.07 | GB (lower is better) |
| vLLM | 8.24 | 1.07 | GB (lower is better) |
| SGLang | 8.91 | 1.07 | GB (lower is better) |
| Ollama + basal-serve | 12.09 | 1.07 | GB (lower is better) |
| llama.cpp + basal-serve | 8.15 | 1.07 | GB (lower is better) |
basal-rs, 1.07 GB
Binaries: 47 MB. The rest is NVIDIA CUDA 12.9 libraries (cudart, cuBLAS, cuBLASLt, cuRAND), totaling 1.02 GB. The basal setup command downloads them only if they are missing from the system.
upstream with PyTorch, 7.08 GB
NVIDIA libraries installed as Python packages: 4.5 GB, torch: 1.6 GB, triton: 0.67 GB, plus the transformers library and HTTP server.
Upstream installed following the Basal v1.5.0 instructions in a clean Python 3.12 environment, excluding Python itself (0.1 GB). Ollama 0.20.3 from the official installer. llama.cpp b11535 from the official CUDA 12.8 package: llama-server, its libraries and cudart. File sizes on disk measured on October 9, 2026.
Short questions and long documents together
Realistic traffic: short questions arrive alongside documents of up to 16k A piece of text, usually part of a word. The model reads text as a sequence of tokens.. p50 is the typical response time; p95 is the time within which 95% of requests finish. High p95 with 32 clients comes from long documents waiting in the queue; short questions have a separate queue and do not wait behind them.
| GPU | 32 clients | p50 / p95 | Sequential | p50 / p95 |
|---|---|---|---|---|
| H100 80 GB | 454 | 1.35 / 12.1 | 381 | 0.03 / 0.87 |
| A100 80 GB | 231 | 2.05 / 22.8 | 197 | 0.07 / 1.68 |
| L40S | 214 | 2.23 / 26.2 | 184 | 0.08 / 1.77 |
| RTX 6000 Ada | 166 | 3.03 / 35.6 | 145 | 0.12 / 2.22 |
| RTX A6000 | 121 | 3.55 / 46.5 | 105 | 0.13 / 3.18 |
basal-1.5-mini accepts shorter documents, up to about 4k tokens, so its longest test documents are shorter.
The longer the document, the bigger the difference
For short documents up to about 2k tokens, upstream is faster on H100. From about 4k tokens, basal-rs is 1.75–2.4 times faster; with five questions about the same document, it is 10–15 times faster because the document is computed once.
Chart data table
| Measurement | upstream v1.5.0, BF16 (fast) | basal-rs, f16 | Unit |
|---|---|---|---|
| 512 tokens | 23 | 31 | ms (lower is better) |
| 1,792 tokens | 84 | 100 | ms (lower is better) |
| 4,096 tokens | 487 | 267 | ms (lower is better) |
| 16,384 tokens | 4,200 | 1,743 | ms (lower is better) |
| 16,384 tokens, 5 questions | 29.5 | 2 | s (lower is better) |
Time for a complete request with one question; the longer the document, the greater basal-rs's advantage. Several questions about the same document share its computation: on RTX 6000 Ada, a 16k-token document with five questions takes 112 s in upstream and 5.3 s in basal-rs.
Basal models on MacBooks
On Macs, basal-rs computes with Apple's interface for GPU computation on Macs.. The model author's server uses Apple's machine-learning library for Macs. in A 16-bit number format with a wider range and lower precision than f16.. The switch shows two comparisons: f16 isolates the difference between engines, while bf16 shows the difference an upstream user would experience. The laptop was plugged in during measurement. Times increase on battery or after it heats up.
Chart data table
| Measurement | upstream v1.5.0, MLX f16 | basal-rs, f16 | Unit |
|---|---|---|---|
| basal-1.5-max | 886 | 526 | ms (lower is better) |
| basal-1.5-4.5B | 372 | 235 | ms (lower is better) |
| basal-1.5-mini | 130 | 80 | ms (lower is better) |
Times from multiple runs: we show upstream's best result and basal-rs's worst. On M2 Max, basal-rs computes 1.9–2.0 times as many decisions per second as upstream in f16.
The same decisions as in full precision
900 questions from nine public datasets, 100 from each. We compare decisions with full-precision FP32 calculations. The only basal-rs f16 differences occur on questions where even FP32 is nearly tied, such as 0.503 versus 0.497.
| Variant | Decisions matching FP32 | Max. difference in The model's raw score for an option, used to calculate its probability. | Max. Total variation: half the sum of absolute probability differences across all options. 0 means the distributions are identical. distance |
|---|---|---|---|
| basal-rs f16 (default) | 900/900 | 0.114 | 0.0092 |
| basal-rs f32 | 900/900 | 0.0003 | 0.00001 |
| upstream BF16 (fast) | 896/900 | 1.210 | 0.052 |
Question datasets
- polemo2-in (KLEJ)
- allegro-reviews (KLEJ)
- cdsc-e (KLEJ)
- dyk (KLEJ)
- cbd (KLEJ)
- ag-news
- emotion
- boolq
- sst5
Questions with dozens of options
Basal models choose from at most 10 options. When there are more, basal-rs divides them into groups, selects the best from each and resolves them in a final round. The model author's server does not do this, so we compare accuracy with the TypeSafe API on datasets with known correct answers.
Chart data table
| Measurement | TypeSafe jev-1.13.0 | basal-1.5-max in basal-rs | Unit |
|---|---|---|---|
| banking77 (77 options, EN) | 159 | 156 | correct (higher is better) |
| clinc150 (150 options, EN) | 185 | 182 | correct (higher is better) |
| massive-pl (59 options, PL) | 159 | 157 | correct (higher is better) |
For basal-rs, one question with dozens of options becomes 14–32 smaller model queries. Measured on RTX 6000 Ada with a 200 W power limit.
250 W or 300 W
We measure basal-rs with basal-1.5-max on an RTX 6000 Ada, whose factory power limit is 300 W. Lowering it to 250 W reduces server throughput by about 20%, while energy per decision stays the same.
| Measurement | 250 W | 300 W (default) |
|---|---|---|
| HTTP server, 1 client | 13.9 requests/s, p50 72 ms | 17.3 requests/s, p50 56 ms |
| HTTP server, 32 clients | 21.8 requests/s | 26.3 requests/s |
| Energy per decision, 32 clients | 11.5 J | 11.4 J |
| Mixed load, 32 clients | 134 requests/min | 161 requests/min |
| 16k-token document, 5 questions | 5.9 s | 5.3 s |
Check compatibility yourself
All you need is basal-rs installed (the basal command) and a copy of its repository. The repository contains full-precision upstream FP32 results for three models. Commands that save results refuse to overwrite an existing path.
# full-precision upstream FP32 results are in the repository
git clone https://github.com/itsoltech/basal-rs && cd basal-rs
# checks model prompts token for token, without a GPU; the model is downloaded to the Hugging Face cache
basal check-prompts --model 4.5B --reference reports/reference-basal-1.5-4.5B-fp32
# the same questions computed by basal-rs, saved in the same format
basal export --model 4.5B --inputs reports/reference-basal-1.5-4.5B-fp32 --out export-f16
# compares tokens and decisions, measures logit and probability differences
basal compare --a reports/reference-basal-1.5-4.5B-fp32 --b export-f16 --out compare.json