Basal models read text and answer questions about it, such as which department should handle a support ticket. They are created by Remek Kinas. We at IT SOL develop basal-rs: an open-source Rust engine that runs these models on NVIDIA GPUs and Apple Silicon Macs. It returns the same decisions as the model author's reference calculations. It is free under the Apache 2.0 license.
Installation target
macOS
$ brew install itsoltech/tap/basal-rs$ basal doctor # shows what is installed and what is missing$ basal serve # http://127.0.0.1:8000
Illustration of the API response format. Values are examples.
900/900decisions matching the model author's full-precision FP32 reference (basal-1.5-max, f16)
0differences in responses to the same request, even under load
1.5–2.8×more requests on the same NVIDIA GPU than with the model author's server (upstream)
1.07 GBfull Linux installation including CUDA libraries (upstream Python environment: at least 7.08 GB)
Get started
From installation to your first decision
Choose your system, start the server and send your first question. Two or three commands are enough on Mac and Linux.
terminal
$ brew install itsoltech/tap/basal-rs$ basal doctor # shows what is installed and what is missing$ basal serve # http://127.0.0.1:8000
Requires an Apple Silicon Mac (tested on M1 Pro and M2 Max). On first startup, the server downloads basal-1.5-4.5B. Update with brew upgrade basal-rs.
First request
request
curl--fail-with-body localhost:8000/v1/systemone \
-H'content-type: application/json'-d '{"model":"basal-1.5-4.5B","state":"Support ticket: a customer requests a refund for a damaged phone.","questions":{"dept":{"type":"choice","instructions":"Which department should handle this ticket?","criteria":{"returns":"Returns","tech":"Technical support","sales":"Sales"}},"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
One basal-rs process serves multiple models at once. No separate server, port or container is needed per model: models, address and limits are defined in basal-serve.yml.
The request selects the model
The model field selects a model, for example basal-1.5-mini for simple questions and basal-1.5-max for harder ones. Requests without this field use the default model. An unknown name returns error 422 with a list of available models.
Shared GPU
Each model has its own queue. The GPU alternates between models, and each model's short batches go ahead of long documents. The mini, 4.5B and max weights together occupy about 36 GB of GPU memory.
Pinned model versions
At startup, the server downloads models from Hugging Face at a pinned commit (revision). If Hugging Face is unavailable, it starts with cached models.
basal-serve.yml
addr: 0.0.0.0:8000default_model: basal-1.5-4.5B # when "model" is omittedmax_inflight: 1024# excess requests receive 529metrics: true # GET /metricsmodels:
- repo: Remek/basal-1.5-mini
- repo: Remek/basal-1.5-4.5B
- repo: Remek/basal-1.5-max
revision: be1b5ee7e7a9755a931262fa7fab4f59be0fd03c
long_tokens: 4096# longer requests use a separate lane
metrics: true or --metrics enables GET /metrics: request counts, response times, tokens and questions, queues and batches, per model. Every response also includes queue-time and compute-time headers.
max_inflight limits the total number of queued and running requests across all models (default: 1024). Additional requests immediately receive status 529 with Retry-After, so the client knows when to try again.
basal doctor checks the system, GPU, driver and CUDA libraries, configuration file, cached models, free GPU memory and disk space, Hugging Face access and port. For each problem, it suggests a fix.
basal setup --service installs the server as a user service: systemd on Linux, launchd on macOS. basal update downloads the latest release. Once a day, the server checks for a newer release and logs the result.
ghcr.io/itsoltech/basal-rs with release tags such as v0.1 or latest. One image runs on A100, RTX 30xx/40xx, L40S and H100. Downloaded models stay in a volume, so subsequent starts take a few seconds. docker stop shuts the server down gracefully.
basal benchmark starts the server and measures it with 1, 4, 8, 16 and 32 clients: p50 and p99 response times, requests per second, RAM and GPU memory. The binary includes 900 test questions; --requests replays requests from your own application.
basal client is a command-line program. It reads data from standard input (stdin) or a file and returns Basal model decisions. Store the result in a Bash variable, pipe it to jq or write it to another JSONL file. With --local, the model runs inside the client: data stays on your computer and no server is needed.
command
cat ticket.txt | basal client --local \
--ask'Does the ticket describe an outage that prevents work?'--value
example output
0.9937360260087164
ticket.txt contains: “After the update, the system will not start and the whole team is unable to work.” --ask alone asks a yes/no question. --value prints only the probability of “yes”, ready to store in a Bash variable or use in a condition.
command
department=$(basal client --local--state complaint.txt \
--ask'Which department should handle this ticket?' \
--choice'returns=Returns and complaints' \
--choice'support=Technical support' \
--choice'sales=Pre-purchase questions' \
--choice'other=Other issue or insufficient information'--value)
echo"$department"
example output
returns
complaint.txt contains: “The phone arrived damaged. I want a refund.” In each --choice, the key before = is returned to the script, while the description after it explains the option to the model. A choice question always returns one of the given options.
Each line of a JSONL file is a separate JSON record. The routing.json template asks two questions about each record: department and urgency. The .input field holds the complete record, so the ID is preserved. In the original run, tickets 103 (a pre-purchase question about a monitor) and 104 were routed to other because the template says to choose other when information is missing. The full response includes the probability of each option, so a threshold or more precise criteria can catch such cases.
command
cat app.log | basal client --local--input lines \
--ask'Does this log entry indicate data loss or a service outage?' |
jq-c'{p: (.response.answers.result.noul * 100 | round / 100), line: .input}'
example output
{"p":0.01,"line":"10:02:11 INFO api: GET /v1/orders 200 in 41 ms"}{"p":0.03,"line":"10:02:13 WARN cache: hit ratio dropped to 71%"}{"p":0.87,"line":"10:02:19 ERROR db: primary unreachable, failing over to replica"}{"p":0.86,"line":"10:02:24 ERROR storage: disk full, 312 uploads not saved"}{"p":0.01,"line":"10:02:31 WARN auth: 3 failed logins for one user"}{"p":0.01,"line":"10:02:40 INFO worker: job 8812 finished in 2.4 s"}{"p":0.92,"line":"10:02:44 ERROR api: 503 for all /v1/payments requests"}
--input lines evaluates each non-empty line separately. The model loads once for the entire stream. For alerts, add select(.p >= threshold) in jq. Set the threshold according to your own rules; it does not guarantee accuracy. For a continuous log stream, replace cat with tail -F.
One template asks five questions about the same state: choice (one option), noul (yes/no), score (weighted average of levels 0 to 3), multi (multiple options) and act (action selection). The state is an object containing refund conditions and the complaint text. act selects the action with the lowest expected cost based on the template's costs. The client returns the decision but does not execute it.
Commands and inline input examples are shown in English. Output values illustrate the format and come from the original Polish runs on an Apple Silicon Mac (M1 Pro), using --local and basal-1.5-4.5B; English results may differ. The original input files and templates are in examples/client in the basal-rs repository. Translate their prompts and input text when using them with the English examples.
Locally or over HTTP
With --local, the client loads the model once for the whole input and works without a server. Otherwise it sends requests to basal serve, up to 256 concurrently (--jobs).
The record comes back with the answer
In streaming mode, each output line includes .input with the entire original record. IDs and other fields are preserved.
Four input formats
text, json, lines and jsonl. --state-pointer specifies which record field the model should evaluate.
Exit codes for scripts
0 when all records finish without errors, including “no” answers. 1 for a record error. 2 for configuration, input or output errors.
We use Basal ourselves. We share the engine with everyone.
We use Basal models in our own projects. We needed an engine that responds in tens of milliseconds on a regular workstation and returns the same decisions as the model author's reference calculations. We wrote it in Rust.
We share the same engine for free, together with its source code and measurements. We want everyone with an Apple Silicon Mac or an NVIDIA GPU to make the most of their hardware. This is how we support Polish AI and make Polish models easier to use.
The Apache 2.0 license allows commercial use, modification and integration into your own products.
02
Everything is public
Code, measurement reports and raw data are available in the public GitHub repository.
03
Compute decisions on your own hardware
The server runs on your computer and accepts only local connections by default (127.0.0.1). It downloads model weights from Hugging Face.
04
The model author is always credited
We credit Remek Kinas as the model author in the README, NOTICE file and on this website.
Benefits
What matters to us
The engine changes how computation is carried out. The model itself stays the same. We verify each claim below with measurements, and publish the reports and data in the public repository.
01 / Simplicity
Two commands to get a server running
One binary, without Python or virtual environments. basal doctor checks your computer and explains what is missing and how to fix it. On Linux, the full installation with CUDA libraries occupies 1.07 GB. The Python environment for the basal-serve, the model author's reference server, written in Python (PyTorch, or MLX on Macs). occupies at least 7.08 GB.
The same decisions as the model author's reference
The author evaluated the models using Numbers stored in 32 bits: full-precision computation.. basal-rs defaults to Numbers stored in 16 bits, half as many bits as FP32. with f32 accumulation. Across 900 questions from nine datasets, its decisions match FP32 calculations (basal-1.5-max, A100).
The same request returns a bitwise-identical response whether it arrives alone, in a batch with others or with 32 concurrent clients. A decision does not depend on what else the server is computing at the time.
Compared with upstream in fast mode: the same GPU and the same precision class, 16 bits (f16 and BF16). basal-rs reduces energy per decision by a factor of 1.3–2.8.
Runs on Apple Silicon Macs and NVIDIA GPUs: A100, RTX 30xx and newer. The model comes in three sizes. Questions can be in Polish or English.
basal-1.5-mini~3.2 GB
basal-1.5-4.5B~9.5 GB
basal-1.5-max~23 GB
06 / Access
A decision in one HTTP request
The server supports the TypeSafe System One API and upstream server conventions. Applications built for either work without changes. One server process can serve multiple models.
POST /v1/systemoneTypeSafe System One contract
POST /v1/basalupstream server conventions
GET /v1/modelslist of available models
GET /metricsoptional Prometheus metrics
Performance
Measurements on the same GPUs
Selected measurements of basal-rs and the basal-serve v1.5.0, the model author's reference server in Python (PyTorch, or MLX on Macs). NVIDIA fast mode uses BF16, torch.compile and CUDA graphs. in its fast mode on the same GPUs and Macs. Both engines use 16-bit precision. All models, GPUs and measurements are listed on the benchmarks page.
basal-1.5-max on RTX 6000 AdaEach row has its own scale and unit.
upstream v1.5.0, BF16 (fast)basal-rs, f16
Single decision, median ms, lower is better
90.7
63.4–64.0
1.4× faster
One question over HTTP, median (p50) ms, lower is better
118
72
1.6× faster
32 concurrent clients requests/s, higher is better
6.8
21.8
3.2× more
Mixed load, 32 clients requests/min, higher is better
17
119
7× more
16k-token document, 5 questions s, lower is better
112.2
5.3
21× faster
Chart data table
Measurement
upstream v1.5.0, BF16 (fast)
basal-rs, f16
Unit
Single decision, median
90.7
63.4–64.0
ms (lower is better)
One question over HTTP, median (p50)
118
72
ms (lower is better)
32 concurrent clients
6.8
21.8
requests/s (higher is better)
Mixed load, 32 clients
17
119
requests/min (higher is better)
16k-token document, 5 questions
112.2
5.3
s (lower is better)
Upstream ran on the same GPU in its fastest mode (fast). GPU power limit: 250 W, or 300 W for the 16k-token document. A token is a piece of text: the model reads input as a sequence of tokens.
basal-rs receives an HTTP request, builds prompts following the Basal protocol, computes them on a GPU or Mac, and returns a probability distribution over the options. The number's color indicates whose code handles that step.
basal-rs engineBasal protocol and models · Remek Kinas
Each question is sent to the model twice, with options in two orders. The result is averaged and calibrated. basal-rs reproduces Remek Kinas's protocol without changes.
Yes. The engine uses the Apache 2.0 license, which also permits commercial use. The Basal model author released the models under Apache 2.0 as well. No account or API key is required. There are no limits.
The Basal models and their author, Remek Kinas. basal-rs reproduces his decision protocol: the same prompt, token IDs, both option orders and calibration. Our measurements describe engine performance and compatibility with upstream reference calculations. They do not evaluate model accuracy.
f16 stores numbers in 16 bits; FP32 uses 32 bits (full precision). Our reference is the FP32 computation used by the author to evaluate the models. basal-rs defaults to f16 with f32 accumulation, matching FP32 decisions on 899–900 of 900 questions. The only differences occur in near ties, such as 0.503 versus 0.497 in FP32. Upstream's fast mode uses BF16, another 16-bit format, and matches on 890–896 of 900 questions. We compare performance against that mode: 16 bits against 16 bits. For FP32 computation, set dtype: f32.
An Apple Silicon Mac, or an x86_64 Linux computer with an NVIDIA GPU with compute capability 8.0 or higher (A100, RTX 30xx or newer). GPU memory must hold model weights plus activations (intermediate results): basal-1.5-mini ~3.2 GB, 4.5B ~9.5 GB, max ~23 GB.
Yes. POST /v1/systemone follows the TypeSafe System One contract, including its error statuses. POST /v1/basal follows the upstream v1.5.0 server conventions. The application does not need to know that basal-rs is on the other end.
The decision protocol is the same; the computation differs. basal-rs arranges prompts into a prefix tree, computing a shared state only once. It has custom CUDA and Metal GPU code and a scheduler that prioritizes short requests over long ones. No Python is needed at runtime.
The server has no authentication and does not check who sends requests. By default, it accepts only local connections (127.0.0.1). The 0.0.0.0 address accepts network connections and is also used in the Docker image. It is intended for a trusted network or for use behind an authenticating proxy.
Roadmap
What works and what comes next
Version 1.0.0 aims to refine the single server: supported systems, GPUs, performance and integrations. Later, the engine is planned to support multiple GPUs and machines. Releases and changelogs are published on GitHub. Issues and pull requests are welcome for every item.
Released (0.1.x)
NVIDIA GPUs (CUDA) and Apple Silicon (Metal)GPU code for NVIDIA compute capabilities 8.0, 8.9 and 9.0 in one binary, plus Apple Silicon support.
basal-1.5 and basal-1.0 modelsbasal-1.5 mini, 4.5B and max with multi, act, facts and evidence extensions, plus basal-1.0-4.5B.
One-command installationHomebrew, install.sh script, Docker image, basal doctor and basal update.
Prometheus metricsRequest counts, response times, queues and batches, per model.
1.0.0
All Basal modelsVersions 1.0, 1.5, 1.8 and 2.0. Support for basal-1.0-1.5B will be added. The author has announced basal-1.8 and 2.0; support will follow publication of their weights.
WindowsServer and command-line tool (CLI) on Windows with NVIDIA GPUs, alongside Linux and macOS.
Bash automation clientbasal client: questions from the command line, data from stdin or JSONL files, concurrent requests and a local model without starting a server.Usage examples
PerformanceLower latency where upstream is currently faster: single decisions and states up to about 2k tokens on H100.
More GPUsGPU code and measurements for additional NVIDIA generations beyond the currently supported compute capabilities 8.0, 8.9 and 9.0.
More testsBroader decision-compatibility tests against the reference and server load tests.
API clients and SDKsLibraries for calling the server from application code without manually constructing HTTP requests.
After 1.0.0
GPU clustersMultiple GPUs share a request queue and a set of models.
Distributed computationServers connect and distribute requests among themselves.
Cluster and deployment managementAdd machines and deploy new model versions from one place.
Admin dashboardMachine, model and queue status in the browser.
Encrypted connections (mTLS)Encrypted, mutually authenticated connections between clients, servers and cluster nodes.
basal-rs runs models but does not create them. Basal models make the decisions; responsibility for decision quality lies with the model author, Remek Kinas. The basal-rs engine is responsible for computation speed and for how the computation is carried out.
ownerYou
Your application
Sends a state (text or data to evaluate) and questions about it over HTTP. Receives decisions with probabilities.
HTTP · POST /v1/systemone
ownerIT SOL
basal-rs
The Rust engine and server that run the models. We are responsible for how quickly and predictably decisions are computed.