Benchmarks

Two engines, the same models on the same GPU

basal-rs runs Basal models. We compare its performance with the basal-serve v1.5.0, the Basal model author's reference server in Python and PyTorch. NVIDIA fast mode uses BF16, torch.compile and CUDA graphs. On Macs it uses MLX., the model author's server, on the same GPU. Both engines use 16-bit precision: basal-rs uses f16 and upstream uses BF16. We compare both engines' decisions with full-precision FP32 calculations, the precision used by the author to validate the models. These results describe the engine; decision accuracy comes from the Basal models.

  • 900/900 decisions matching the full-precision FP32 reference (basal-1.5-max, A100, f16)
  • 0 differences in responses to the same request under load
  • 1.5–2.8× more requests on the same GPU (14 of 15 GPU/model pairs)
  • 1.3–2.8× less energy per decision across all 15 GPU/model pairs
01 / Methodology

How we measure

Performance and compatibility results come from reports in the basal-rs repository, with data and test conditions. Both engines are always measured on the same GPU, with basal-rs using its default installation settings.

We compare the performance of basal-rs in f16 with upstream in BF16 on the same GPU. Decisions from both engines are compared with upstream FP32 calculations.
our engine, default settings

basal-rs, f16

f16 weights and intermediate results, f32 accumulation in matrix multiplication

the model author's server, fast mode

upstream, BF16

PyTorch with torch.compile and CUDA graphs; MLX in bf16 on Macs

899–900 / 900 decisions matching FP32
890–896 / 900 decisions matching FP32
reference for decision quality

upstream, FP32

The full precision used by the author to evaluate the models. Used to compare decisions, not performance. For FP32 computation in basal-rs, set dtype: f32 to get 900 / 900.

  • Single decision

    Time for one decision when the server has no other work: 44 sample questions, each in both option orders. Median of 39 measurements after discarding 5.

  • HTTP server

    A running server with 1, 8 or 32 concurrent clients; each request contains one question. The same program sends requests to both servers, and we read energy usage from the GPU.

  • Mixed load

    Documents from 100 to 16k tokens and 1 to 14 questions per request. Only basal-rs was measured: upstream handles about 17 such requests per minute on the max model, so measuring it would take over an extra hour per GPU.

  • Compatibility

    We compare decisions with full-precision FP32 calculations, the precision used by the model author for validation. Both engines must construct exactly the same model input, token for token.

GPUs rented from Shadeform, one per machine, at factory power limits. NVIDIA driver 580, or 535 with the CUDA compatibility package on H100. Measured on October 8, 2026, with basal-rs 0.1.4.
GPUArchitecture, Compute capability: the NVIDIA architecture version that determines which instructions a GPU supports.MemoryPower limitNumber of virtual CPU cores (vCPUs) on the host machine.
H100 80 GBHopper, 9.080 GB700 W16 vCPU
A100 80 GBAmpere, 8.080 GB400 W16 vCPU
L40SAda, 8.948 GB350 W12 vCPU
RTX 6000 AdaAda, 8.948 GB300 W12 vCPU
RTX A6000Ampere, 8.648 GB300 W6 vCPU
02 / NVIDIA GPUs

Five GPUs, three models

The controls above the chart change the model and measurement. The number marked × shows how much better the winning engine's result is for each GPU. When upstream wins, its name precedes the number.

HTTP server throughput with 32 clients basal-1.5-max, requests/s, higher is better
upstream v1.5.0, BF16 (fast) basal-rs, f16
Chart data table
Measurementupstream v1.5.0, BF16 (fast)basal-rs, f16Unit
H100 80 GB5685requests/s (higher is better)
A100 80 GB2441requests/s (higher is better)
L40S1836requests/s (higher is better)
RTX 6000 Ada1027requests/s (higher is better)
RTX A60001021requests/s (higher is better)

In 14 of 15 GPU/model pairs, basal-rs handles 1.5–2.8 times as many requests as upstream. The exception is basal-1.5-mini on H100, where the two servers are similar (basal-rs reaches 0.95 times upstream's result).

03 / Installation size

Without Python or PyTorch

The chart shows the disk space required to install everything needed to serve Basal models on an NVIDIA GPU. Model files and the GPU driver are excluded. In every mode, upstream runs as basal-serve in Python with torch and transformers, including when Ollama or llama.cpp performs the computation.

Linux installation with an NVIDIA GPU GB on disk, excluding model files and driver, lower is better
upstream: basal-serve + engine basal-rs 0.1.6
Chart data table
Measurementupstream: basal-serve + enginebasal-rs 0.1.6Unit
PyTorch (fast mode)7.081.07GB (lower is better)
vLLM8.241.07GB (lower is better)
SGLang8.911.07GB (lower is better)
Ollama + basal-serve12.091.07GB (lower is better)
llama.cpp + basal-serve8.151.07GB (lower is better)

basal-rs, 1.07 GB

Binaries: 47 MB. The rest is NVIDIA CUDA 12.9 libraries (cudart, cuBLAS, cuBLASLt, cuRAND), totaling 1.02 GB. The basal setup command downloads them only if they are missing from the system.

upstream with PyTorch, 7.08 GB

NVIDIA libraries installed as Python packages: 4.5 GB, torch: 1.6 GB, triton: 0.67 GB, plus the transformers library and HTTP server.

Upstream installed following the Basal v1.5.0 instructions in a clean Python 3.12 environment, excluding Python itself (0.1 GB). Ollama 0.20.3 from the official installer. llama.cpp b11535 from the official CUDA 12.8 package: llama-server, its libraries and cudart. File sizes on disk measured on October 9, 2026.

04 / Mixed load

Short questions and long documents together

Realistic traffic: short questions arrive alongside documents of up to 16k A piece of text, usually part of a word. The model reads text as a sequence of tokens.. p50 is the typical response time; p95 is the time within which 95% of requests finish. High p95 with 32 clients comes from long documents waiting in the queue; short questions have a separate queue and do not wait behind them.

basal-rs only: requests handled per minute and p50 / p95 response times in seconds. With 32 clients, requests arrive concurrently; in the Sequential column, they arrive one at a time.
GPU32 clientsp50 / p95Sequentialp50 / p95
H100 80 GB 4541.35 / 12.13810.03 / 0.87
A100 80 GB 2312.05 / 22.81970.07 / 1.68
L40S 2142.23 / 26.21840.08 / 1.77
RTX 6000 Ada 1663.03 / 35.61450.12 / 2.22
RTX A6000 1213.55 / 46.51050.13 / 3.18

basal-1.5-mini accepts shorter documents, up to about 4k tokens, so its longest test documents are shorter.

05 / Long documents

The longer the document, the bigger the difference

For short documents up to about 2k tokens, upstream is faster on H100. From about 4k tokens, basal-rs is 1.75–2.4 times faster; with five questions about the same document, it is 10–15 times faster because the document is computed once.

Complete request time on H100 PCIe 350 W median, one client, without HTTP; each row has its own scale
upstream v1.5.0, BF16 (fast) basal-rs, f16
Chart data table
Measurementupstream v1.5.0, BF16 (fast)basal-rs, f16Unit
512 tokens2331ms (lower is better)
1,792 tokens84100ms (lower is better)
4,096 tokens487267ms (lower is better)
16,384 tokens4,2001,743ms (lower is better)
16,384 tokens, 5 questions29.52s (lower is better)

Time for a complete request with one question; the longer the document, the greater basal-rs's advantage. Several questions about the same document share its computation: on RTX 6000 Ada, a 16k-token document with five questions takes 112 s in upstream and 5.3 s in basal-rs.

06 / Apple Silicon

Basal models on MacBooks

On Macs, basal-rs computes with Apple's interface for GPU computation on Macs.. The model author's server uses Apple's machine-learning library for Macs. in A 16-bit number format with a wider range and lower precision than f16.. The switch shows two comparisons: f16 isolates the difference between engines, while bf16 shows the difference an upstream user would experience. The laptop was plugged in during measurement. Times increase on battery or after it heats up.

Single-decision latency on M2 Max median in milliseconds, lower is better
upstream v1.5.0, MLX f16 basal-rs, f16
Chart data table
Measurementupstream v1.5.0, MLX f16basal-rs, f16Unit
basal-1.5-max886526ms (lower is better)
basal-1.5-4.5B372235ms (lower is better)
basal-1.5-mini13080ms (lower is better)

Times from multiple runs: we show upstream's best result and basal-rs's worst. On M2 Max, basal-rs computes 1.9–2.0 times as many decisions per second as upstream in f16.

07 / Compatibility

The same decisions as in full precision

900 questions from nine public datasets, 100 from each. We compare decisions with full-precision FP32 calculations. The only basal-rs f16 differences occur on questions where even FP32 is nearly tied, such as 0.503 versus 0.497.

A100 80 GB. Decisions matching upstream full-precision FP32 calculations and the largest differences in model scores. f16, f32 and BF16 are computation precisions.
VariantDecisions matching FP32Max. difference in The model's raw score for an option, used to calculate its probability.Max. Total variation: half the sum of absolute probability differences across all options. 0 means the distributions are identical. distance
basal-rs f16 (default)900/9000.1140.0092
basal-rs f32900/9000.00030.00001
upstream BF16 (fast)896/9001.2100.052

Question datasets

  • polemo2-in (KLEJ) Polish · choice · options: 4
  • allegro-reviews (KLEJ) Polish · score · options: 5
  • cdsc-e (KLEJ) Polish · choice · options: 3
  • dyk (KLEJ) Polish · yes/no · options: 2
  • cbd (KLEJ) Polish · yes/no · options: 2
  • ag-news English · choice · options: 4
  • emotion English · choice · options: 6
  • boolq English · yes/no · options: 2
  • sst5 English · score · options: 5
08 / Questions with many options

Questions with dozens of options

Basal models choose from at most 10 options. When there are more, basal-rs divides them into groups, selects the best from each and resolves them in a final round. The model author's server does not do this, so we compare accuracy with the TypeSafe API on datasets with known correct answers.

Correct answers on 200 questions choice questions with 59–150 options, higher is better
TypeSafe jev-1.13.0 basal-1.5-max in basal-rs
Chart data table
MeasurementTypeSafe jev-1.13.0basal-1.5-max in basal-rsUnit
banking77 (77 options, EN)159156correct (higher is better)
clinc150 (150 options, EN)185182correct (higher is better)
massive-pl (59 options, PL)159157correct (higher is better)

For basal-rs, one question with dozens of options becomes 14–32 smaller model queries. Measured on RTX 6000 Ada with a 200 W power limit.

09 / Power limit

250 W or 300 W

We measure basal-rs with basal-1.5-max on an RTX 6000 Ada, whose factory power limit is 300 W. Lowering it to 250 W reduces server throughput by about 20%, while energy per decision stays the same.

Measurement250 W300 W (default)
HTTP server, 1 client13.9 requests/s, p50 72 ms17.3 requests/s, p50 56 ms
HTTP server, 32 clients21.8 requests/s26.3 requests/s
Energy per decision, 32 clients11.5 J11.4 J
Mixed load, 32 clients134 requests/min161 requests/min
16k-token document, 5 questions5.9 s5.3 s
10 / Reproduce

Check compatibility yourself

All you need is basal-rs installed (the basal command) and a copy of its repository. The repository contains full-precision upstream FP32 results for three models. Commands that save results refuse to overwrite an existing path.

terminal
# full-precision upstream FP32 results are in the repository
git clone https://github.com/itsoltech/basal-rs && cd basal-rs

# checks model prompts token for token, without a GPU; the model is downloaded to the Hugging Face cache
basal check-prompts --model 4.5B --reference reports/reference-basal-1.5-4.5B-fp32

# the same questions computed by basal-rs, saved in the same format
basal export --model 4.5B --inputs reports/reference-basal-1.5-4.5B-fp32 --out export-f16

# compares tokens and decisions, measures logit and probability differences
basal compare --a reports/reference-basal-1.5-4.5B-fp32 --b export-f16 --out compare.json