From an HTTP request to a response with probabilities, step by step. Remek Kinas developed the Basal decision models and their protocol: the rules for asking questions and computing decisions. The basal-rs engine handles request batching, scheduling, Programs that run directly on the GPU and perform the model's computations. and the server.
basal-rs engineBasal protocol and models · Remek Kinas
01 / Request path
From request to decision
Every request passes through six stages. The stage color identifies whose code handles it: yellow for the basal-rs engine, lavender for the Basal protocol.
Stage 1 of 6
basal-rs engine
Receiving the request
The HTTP server (axum and tokio libraries in Rust) accepts a POST request at /v1/systemone or /v1/basal. The model field selects the model, and the server passes the request to that model's thread.
when more than 1024 requests are in flight (the max_inflight limit), the server returns status 529 with Retry-After: 1, asking the client to retry in one second
each model has its own thread, queue and long-request lane
requests that the client is no longer waiting for are skipped
basal-rs engineBasal protocol · Remek Kinas
Validation and planning
The model thread validates the request, tokenizes the text and builds prompts according to the Basal contract: two prompts per question, with options in two orders. Invalid requests are rejected immediately, before entering the queue.
the prompt template, language detection rule and option letters match upstream (prompt.py)
request cost is the number of tokens to compute when shared prefixes are counted once
the engine will not load a model with a different prompt contract
basal-rs engine
HRRN queue
The queue uses HRRN (Highest Response Ratio Next). It selects the request with the highest ratio of (waiting time + estimated runtime) / estimated runtime. Short requests go ahead of long ones, while long requests move up the queue as they wait.
requests over 4096 tokens go to a separate long-request lane
A batch is a group of requests computed together on the GPU: the first queued request and each subsequent one that fits within 8192 tokens. All prompts in the batch are arranged into rows. Each row is a tree whose shared prompt prefixes are computed once.
a row holds up to 32,768 tokens; one forward pass processes up to 8192
tree depth of up to 8 blocks
identical prompts share a single output readout
basal-rs engine
GPU forward pass
The Llama model processes the entire batch in f16 as one compact token list, without padding to equalize prompt lengths. In the final layer, the MLP block and logits (raw scores before conversion to probabilities) are computed only for readout positions and answer letters.
matrix multiplication (GEMM): cuBLASLt with an algorithm table on CUDA, or code ported from MLX on Metal
f32 attention over tree blocks, using custom CUDA and Metal kernels
f32 accumulation in matrix multiplication
Basal protocol · Remek Kinasbasal-rs engine
Decision and response
Letter logits become probabilities for both option orders. After restoring the request's original option order, the two distributions are averaged and calibrated using the temperature assigned to the question type in CALIBRATION.json. The server builds a TypeSafe System One API response, with confidence calculated using TypeSafe's formulas.
response headers: x-basal-queue-ms (queue time), x-basal-compute-ms (compute time) and x-basal-batch-requests (requests in the batch)
full-precision probabilities (TypeSafe rounds them to two decimal places)
02 / Basal protocol
Two option orders and calibration
The model answers with an option letter. A decision is a probability distribution over those letters, calculated from their The model's raw score for one letter. Softmax converts all letter logits into probabilities that sum to 1.. Each question is sent to the model twice, with options in two orders. Averaging reduces the effect of option order on the result.
Remek Kinas developed this protocol. basal-rs reproduces it without changes: the same The text sent to the model: template, state, question and options., the same A token is a piece of text, usually part of a word. The model receives text as a sequence of token IDs., both orders and calibration from CALIBRATION.json. The second formula below the example describes calibration: the logarithm of the mean is divided by temperature T for that question type, then passed through softmax again.
1
Order 1
options in the request's original order
A Returns0.89
logit 3.2
B Technical support0.09
logit 0.9
C Sales0.02
logit -0.4
2
Order 2
the same options in reverse order
A Sales0.08
logit 0.4
B Technical support0.17
logit 1.1
C Returns0.75
logit 2.6
3
Average
after restoring the original option order
Returns0.82
Technical support0.13
Sales0.05
4
After calibration
temperature T = 1.00
Returns0.82
Technical support0.13
Sales0.05
T < 1 sharpens the distribution; T > 1 flattens it. The selected option stays the same; confidence changes.
Questions about the same document share the beginning of their prompts: a common The beginning of a prompt, shared by multiple prompts.. The engine arranges prompts from one A group of requests computed together on the GPU. into a tree and computes each shared prefix once. The shared parts are the template, The context being asked about, such as a document. and the question text, which is the same in both orders.
Questions about the same state
State length
Each prompt separately1,2901,290 tokens
templatestatequestionoptions
Q1 · order 1
Q1 · order 2
Q2 · order 1
Q2 · order 2
Q3 · order 1
Q3 · order 2
Q4 · order 1
Q4 · order 2
Number of prompts:
6, each with
215
tokens. The state is computed
6
times.
Shared prefix tree300300 tokens
template · 53 · computed once at startup
state · 120
question 1 · 24
order 1 · 18
order 2 · 18
question 2 · 24
order 1 · 18
order 2 · 18
question 3 · 24
order 1 · 18
order 2 · 18
question 4 · 24
order 1 · 18
order 2 · 18
4.3× fewer tokens to compute. The state is computed once for all questions.
Numbers next to blocks are token counts. State, question and option lengths are examples. The Polish prompt template has 53 tokens.
Attention over tree blocks
Before computation, the tree becomes one row of tokens. In The part of the model where each token uses information from earlier tokens. a token sees earlier tokens in its own block and ancestor blocks. A token's position for A way of encoding token positions so the model knows the text order. is its position in its own prompt. Click a row to see which blocks it can attend to.
Packing prompts into a prefix tree comes from The reference basal-serve server by the Basal model author. 1.5 (engine.py). basal-rs also combines questions from different requests in the same batch into one tree. Custom CUDA and Metal kernels compute attention over the tree blocks.
queries ↓ / keys →SQ1Q1·1Q1·2Q2Q2·1Q2·2
Q1 · order 2 attends to:
state, question 1, Q1 · order 2. Sibling blocks (other branches of the same parent) are masked, so the result matches computing each full prompt separately.
04 / Scheduling
Short requests do not wait behind long ones
A 16k-token document takes several seconds to process. To avoid holding up short requests, requests over 4096 tokens enter a separate lane. That lane shares the GPU with short batches and yields it between The model computes its result in stages, layer by layer. of the model.
16k-token document
short requests
time
A short request waits
until the entire document is processeduntil the document's current layer finishes
The document finishes
earliest, because it runs uninterruptedlater by the time spent on short batches
document processingbatch of short requestswaiting in the queuerequest arrival
Illustrative diagram, not a measurement. FIFO processes requests in arrival order. The long-request lane yields the GPU after a layer if it has run for at least long_slice_ms (default: 100 ms). It then resumes that layer with the same tensors (intermediate model data).
priority = (waiting time + estimated runtime) / estimated runtimeHighest Response Ratio Next: the queue selects the request with the highest priority according to this formula.: short requests go first, while long requests move up the queue as they wait.
max_batch_tokens: 8192Batch token limit. A batch consists of the first queued request and each subsequent request that fits.
long_tokens: 4096Long-request lane threshold: requests above this length go to the long-request lane. Set to 0 to disable the lane.
multiple models, one GPUWhen multiple models share a GPU, short batches from every model go ahead of all long requests.
05 / Reproducibility
The same response regardless of traffic
On a GPU, matrix multiplication results can depend on batch row count because the library may choose a different algorithm. basal-rs fixes its algorithms so a question's result does not depend on which other requests share the batch.
alone
in a batch with other requests
in a tree with other questions
under a load of 32 clients
→
Batch-independent GEMM table
General matrix multiplication. is computed by NVIDIA's cuBLASLt library. For each weight shape, the table specifies the fastest algorithm without A matrix multiplication method that combines partial results at the end., selected from algorithms that produce bitwise-identical results. Attention is also tiled by In attention, a token (query) compares itself with earlier tokens (keys). position, and answer letters are read out using a custom kernel.
→
Raw model scores for letters before they are converted into probabilities. are bitwise identical
basal-rs: response difference for the same request
0
upstream: difference depending on the batch
0.001–0.068
single-decision time with the table / with automatic cuBLASLt algorithm selection (max, RTX 6000 Ada)
63.1 / 68.5 ms
On Macs (Metal), matrix multiplication (GEMM) uses code ported from MLX, whose result is independent of row count. Tree-block attention produces the same result alone, in a batch and in a tree. On NVIDIA GPUs, the GEMM table depends on the GPU, model, cuBLASLt version and precision. The server generates it at first startup or uses a table embedded in the binary.
A backend is the part of the engine that computes on specific hardware: CUDA for NVIDIA GPUs, Metal for Apple Silicon Macs. Both run the Llama model's One pass through the model, from input tokens to output scores. pass using the Rust candle library and custom kernels. The default The number format used for computation. f16 and bf16 store numbers in 16 bits; f32 uses 32 bits, providing more precision at a lower speed. is f16, with f32 accumulation in matrix multiplication. f32 precision serves as the reference.
CUDA
NVIDIA GPUs with compute capability 8.0 or later: A100, RTX 30xx/40xx, RTX 6000 Ada, L40S, H100
matrix multiplication (GEMM) via NVIDIA's cuBLASLt library, with an algorithm table selected by exhaustive search (basal gemm-search)
attention on tensor cores, the GPU's matrix multiplication units; a separate H100 kernel uses its wgmma instructions and produces bitwise-identical results
one binary with kernels for compute capabilities 8.0, 8.9 and 9.0; the right kernels are selected at startup
precomputed GEMM tables for H100 and RTX 6000 Ada embedded in the binary
Metal
Apple Silicon Macs, tested on M1 Pro and M2 Max (32 GB)
matrix multiplication (GEMM) ported from MLX, with results independent of row count
f32 attention on 8×8 matrix blocks (Metal's simdgroup_float8x8 type)
no Python required at runtime
Compute precision and decision agreement with upstream FP32 on 900 questions
Precision
Decisions matching FP32
Computation
f16default
899–900 / 900
bf16 weights converted to f16 on loading; residual stream, normalizations and MLP blocks in f16; matrix multiplication accumulation in f32
bf16
not measured
the same precision as the stored model weights, with no conversion
f32reference
900 / 900
bf16 weights converted to f32 before use; logits differ from upstream FP32 by at most 0.0003; slower computation
07 / Code
Three Rust crates
basal-coreSystem One API contract, prompt construction, tokenizer, prefix-tree packing, decisions and confidence, facts and evidence extensions, choice questions with 11–255 options, engine with a shared backend interface (the Backend trait), GPU sharing between lanes
basal-gpuloading safetensors weights, Llama forward pass using candle, CUDA and Metal kernels, matrix multiplication via cuBLASLt or code ported from MLX, tree-block attention, evidence extension head
basal-clithe basal terminal command, HTTP server, exports and comparisons with the upstream reference, measurements