Open-source model inference engine

Run Basal decision models on your hardware

Basal models read text and answer questions about it, such as which department should handle a support ticket. They are created by Remek Kinas. We at IT SOL develop basal-rs: an open-source Rust engine that runs these models on NVIDIA GPUs and Apple Silicon Macs. It returns the same decisions as the model author's reference calculations. It is free under the Apache 2.0 license.

Installation target
macOS
brew install itsoltech/tap/basal-rsbasal doctor  # shows what is installed and what is missingbasal serve  # http://127.0.0.1:8000
Requirements and next steps
POST /v1/systemone
state

Support ticket: a customer requests a refund for a damaged phone.

200 OK
dept choice Which department should handle this ticket?
Returns 0.94
Technical support 0.05
Sales 0.01
choice returns · confidence 0.91
urgent noul Is this urgent?
P(yes) 0.31
state

Comment below an article: “The author has no idea what they are writing about. A typical amateur.”

200 OK
tone score How harsh is the comment's tone?
0 · neutral 0.06
1 · critical 0.71
2 · insulting 0.23
score 1.17 · confidence 0.57
remove noul Does the comment violate the rules?
P(yes) 0.22
state

Email: “Attached is invoice No. 14/10/2026 for air conditioning maintenance. Payment due in 14 days.”

200 OK
doc choice What type of document is this?
Invoice 0.97
Quote 0.01
Complaint 0.01
Other 0.01
choice invoice · confidence 0.96
approve noul Does it require payment approval?
P(yes) 0.88
engine: basal-rs model: basal-1.5-max · Remek Kinas

Illustration of the API response format. Values are examples.

  • 900/900 decisions matching the model author's full-precision FP32 reference (basal-1.5-max, f16)
  • 0 differences in responses to the same request, even under load
  • 1.5–2.8× more requests on the same NVIDIA GPU than with the model author's server (upstream)
  • 1.07 GB full Linux installation including CUDA libraries (upstream Python environment: at least 7.08 GB)
Get started

From installation to your first decision

Choose your system, start the server and send your first question. Two or three commands are enough on Mac and Linux.

terminal
brew install itsoltech/tap/basal-rsbasal doctor  # shows what is installed and what is missingbasal serve  # http://127.0.0.1:8000

Requires an Apple Silicon Mac (tested on M1 Pro and M2 Max). On first startup, the server downloads basal-1.5-4.5B. Update with brew upgrade basal-rs.

First request

request
curl --fail-with-body localhost:8000/v1/systemone \
  -H 'content-type: application/json' -d '{
  "model": "basal-1.5-4.5B",
  "state": "Support ticket: a customer requests a refund for a damaged phone.",
  "questions": {
    "dept": {"type": "choice", "instructions": "Which department should handle this ticket?",
             "criteria": {"returns": "Returns", "tech": "Technical support", "sales": "Sales"}},
    "urgent": {"type": "noul", "instructions": "Is this urgent?"}
  }
}'
response, example values
{
  "model": "basal-1.5-4.5B",
  "answers": {
    "dept": {
      "type": "choice",
      "choice": "returns",
      "probabilities": {"returns": 0.94, "tech": 0.05, "sales": 0.01},
      "confidence": 0.91
    },
    "urgent": {"type": "noul", "noul": 0.31}
  }
}
Capabilities

One server for all models

One basal-rs process serves multiple models at once. No separate server, port or container is needed per model: models, address and limits are defined in basal-serve.yml.

The request selects the model
The model field selects a model, for example basal-1.5-mini for simple questions and basal-1.5-max for harder ones. Requests without this field use the default model. An unknown name returns error 422 with a list of available models.
Shared GPU
Each model has its own queue. The GPU alternates between models, and each model's short batches go ahead of long documents. The mini, 4.5B and max weights together occupy about 36 GB of GPU memory.
Pinned model versions
At startup, the server downloads models from Hugging Face at a pinned commit (revision). If Hugging Face is unavailable, it starts with cached models.
basal-serve.yml
addr: 0.0.0.0:8000
default_model: basal-1.5-4.5B   # when "model" is omitted
max_inflight: 1024              # excess requests receive 529
metrics: true                   # GET /metrics

models:
  - repo: Remek/basal-1.5-mini
  - repo: Remek/basal-1.5-4.5B
  - repo: Remek/basal-1.5-max
    revision: be1b5ee7e7a9755a931262fa7fab4f59be0fd03c
    long_tokens: 4096           # longer requests use a separate lane
terminal
basal init --model mini --model 4.5B --model maxbasal serve  # reads basal-serve.yml
  • Prometheus metrics

    metrics: true or --metrics enables GET /metrics: request counts, response times, tokens and questions, queues and batches, per model. Every response also includes queue-time and compute-time headers.

    x-basal-queue-ms, x-basal-compute-ms Metrics and queries
  • Limits instead of congestion

    max_inflight limits the total number of queued and running requests across all models (default: 1024). Additional requests immediately receive status 529 with Retry-After, so the client knows when to try again.

    529 Retry-After All configuration options
  • Diagnostics in one command

    basal doctor checks the system, GPU, driver and CUDA libraries, configuration file, cached models, free GPU memory and disk space, Hugging Face access and port. For each problem, it suggests a fix.

    basal doctor --config basal-serve.yml basal commands
  • Service and updates

    basal setup --service installs the server as a user service: systemd on Linux, launchd on macOS. basal update downloads the latest release. Once a day, the server checks for a newer release and logs the result.

    basal setup --service Installation and updates
  • Docker image

    ghcr.io/itsoltech/basal-rs with release tags such as v0.1 or latest. One image runs on A100, RTX 30xx/40xx, L40S and H100. Downloaded models stay in a volume, so subsequent starts take a few seconds. docker stop shuts the server down gracefully.

    docker compose up -d Running with Docker
  • Measure on your own hardware

    basal benchmark starts the server and measures it with 1, 4, 8, 16 and 32 clients: p50 and p99 response times, requests per second, RAM and GPU memory. The binary includes 900 test questions; --requests replays requests from your own application.

    basal benchmark --model mini Choosing a model for your machine
Terminal automation

Decisions in scripts and data pipelines

basal client is a command-line program. It reads data from standard input (stdin) or a file and returns Basal model decisions. Store the result in a Bash variable, pipe it to jq or write it to another JSONL file. With --local, the model runs inside the client: data stays on your computer and no server is needed.

command
cat ticket.txt | basal client --local \
  --ask 'Does the ticket describe an outage that prevents work?' --value
example output
0.9937360260087164

ticket.txt contains: “After the update, the system will not start and the whole team is unable to work.” --ask alone asks a yes/no question. --value prints only the probability of “yes”, ready to store in a Bash variable or use in a condition.

command
department=$(basal client --local --state complaint.txt \
  --ask 'Which department should handle this ticket?' \
  --choice 'returns=Returns and complaints' \
  --choice 'support=Technical support' \
  --choice 'sales=Pre-purchase questions' \
  --choice 'other=Other issue or insufficient information' --value)
echo "$department"
example output
returns

complaint.txt contains: “The phone arrived damaged. I want a refund.” In each --choice, the key before = is returned to the script, while the description after it explains the option to the model. A choice question always returns one of the given options.

command
basal client --local --template routing.json \
  --state tickets.jsonl --input jsonl --state-pointer /text |
  jq -c '{id: .input.id, department: .response.answers.department.choice,
    urgent: (.response.answers.urgent.noul * 100 | round / 100)}'
example output
{"id":101,"department":"returns","urgent":0.03}
{"id":102,"department":"support","urgent":0.02}
{"id":103,"department":"other","urgent":0.06}
{"id":104,"department":"other","urgent":0.98}

Each line of a JSONL file is a separate JSON record. The routing.json template asks two questions about each record: department and urgency. The .input field holds the complete record, so the ID is preserved. In the original run, tickets 103 (a pre-purchase question about a monitor) and 104 were routed to other because the template says to choose other when information is missing. The full response includes the probability of each option, so a threshold or more precise criteria can catch such cases.

command
cat app.log | basal client --local --input lines \
  --ask 'Does this log entry indicate data loss or a service outage?' |
  jq -c '{p: (.response.answers.result.noul * 100 | round / 100), line: .input}'
example output
{"p":0.01,"line":"10:02:11 INFO  api: GET /v1/orders 200 in 41 ms"}
{"p":0.03,"line":"10:02:13 WARN  cache: hit ratio dropped to 71%"}
{"p":0.87,"line":"10:02:19 ERROR db: primary unreachable, failing over to replica"}
{"p":0.86,"line":"10:02:24 ERROR storage: disk full, 312 uploads not saved"}
{"p":0.01,"line":"10:02:31 WARN  auth: 3 failed logins for one user"}
{"p":0.01,"line":"10:02:40 INFO  worker: job 8812 finished in 2.4 s"}
{"p":0.92,"line":"10:02:44 ERROR api: 503 for all /v1/payments requests"}

--input lines evaluates each non-empty line separately. The model loads once for the entire stream. For alerts, add select(.p >= threshold) in jq. Set the threshold according to your own rules; it does not guarantee accuracy. For a continuous log stream, replace cat with tail -F.

command
basal client --local --template all-types.json \
  --state complaint.json --input json |
  jq -c 'def r: . * 100 | round / 100; .answers | {department: .department.choice,
    urgent: (.urgent.noul | r), severity: (.severity.score | r),
    tags: .tags.selected, action: .refund_action.action}'
example output
{"department":"returns","urgent":0.02,"severity":2.11,"tags":["damage","refund","delivery"],"action":"approve"}

One template asks five questions about the same state: choice (one option), noul (yes/no), score (weighted average of levels 0 to 3), multi (multiple options) and act (action selection). The state is an object containing refund conditions and the complaint text. act selects the action with the lowest expected cost based on the template's costs. The client returns the decision but does not execute it.

Commands and inline input examples are shown in English. Output values illustrate the format and come from the original Polish runs on an Apple Silicon Mac (M1 Pro), using --local and basal-1.5-4.5B; English results may differ. The original input files and templates are in examples/client in the basal-rs repository. Translate their prompts and input text when using them with the English examples.

Locally or over HTTP
With --local, the client loads the model once for the whole input and works without a server. Otherwise it sends requests to basal serve, up to 256 concurrently (--jobs).
The record comes back with the answer
In streaming mode, each output line includes .input with the entire original record. IDs and other fields are preserved.
Four input formats
text, json, lines and jsonl. --state-pointer specifies which record field the model should evaluate.
Exit codes for scripts
0 when all records finish without errors, including “no” answers. 1 for a record error. 2 for configuration, input or output errors.
Mission

We use Basal ourselves. We share the engine with everyone.

We use Basal models in our own projects. We needed an engine that responds in tens of milliseconds on a regular workstation and returns the same decisions as the model author's reference calculations. We wrote it in Rust.

We share the same engine for free, together with its source code and measurements. We want everyone with an Apple Silicon Mac or an NVIDIA GPU to make the most of their hardware. This is how we support Polish AI and make Polish models easier to use.

  • 01

    Free, no registration

    The Apache 2.0 license allows commercial use, modification and integration into your own products.

  • 02

    Everything is public

    Code, measurement reports and raw data are available in the public GitHub repository.

  • 03

    Compute decisions on your own hardware

    The server runs on your computer and accepts only local connections by default (127.0.0.1). It downloads model weights from Hugging Face.

  • 04

    The model author is always credited

    We credit Remek Kinas as the model author in the README, NOTICE file and on this website.

Benefits

What matters to us

The engine changes how computation is carried out. The model itself stays the same. We verify each claim below with measurements, and publish the reports and data in the public repository.

01 / Simplicity

Two commands to get a server running

One binary, without Python or virtual environments. basal doctor checks your computer and explains what is missing and how to fix it. On Linux, the full installation with CUDA libraries occupies 1.07 GB. The Python environment for the basal-serve, the model author's reference server, written in Python (PyTorch, or MLX on Macs). occupies at least 7.08 GB.

$ brew install itsoltech/tap/basal-rs
$ basal serve
Size comparison
02 / Model fidelity

900/900

The same decisions as the model author's reference

The author evaluated the models using Numbers stored in 32 bits: full-precision computation.. basal-rs defaults to Numbers stored in 16 bits, half as many bits as FP32. with f32 accumulation. Across 900 questions from nine datasets, its decisions match FP32 calculations (basal-1.5-max, A100).

Compatibility report
03 / Reproducibility

0

response differences under load

The same request returns a bitwise-identical response whether it arrives alone, in a batch with others or with 32 concurrent clients. A decision does not depend on what else the server is computing at the time.

How it works
04 / Hardware performance

1.5–2.8×

more decisions from the same GPU

Compared with upstream in fast mode: the same GPU and the same precision class, 16 bits (f16 and BF16). basal-rs reduces energy per decision by a factor of 1.3–2.8.

Benchmarks
05 / Every user

From a MacBook to an H100 server

Runs on Apple Silicon Macs and NVIDIA GPUs: A100, RTX 30xx and newer. The model comes in three sizes. Questions can be in Polish or English.

  • basal-1.5-mini ~3.2 GB
  • basal-1.5-4.5B ~9.5 GB
  • basal-1.5-max ~23 GB
06 / Access

A decision in one HTTP request

The server supports the TypeSafe System One API and upstream server conventions. Applications built for either work without changes. One server process can serve multiple models.

  • POST /v1/systemoneTypeSafe System One contract
  • POST /v1/basalupstream server conventions
  • GET /v1/modelslist of available models
  • GET /metricsoptional Prometheus metrics
Performance

Measurements on the same GPUs

Selected measurements of basal-rs and the basal-serve v1.5.0, the model author's reference server in Python (PyTorch, or MLX on Macs). NVIDIA fast mode uses BF16, torch.compile and CUDA graphs. in its fast mode on the same GPUs and Macs. Both engines use 16-bit precision. All models, GPUs and measurements are listed on the benchmarks page.

basal-1.5-max on RTX 6000 Ada Each row has its own scale and unit.
upstream v1.5.0, BF16 (fast) basal-rs, f16
Chart data table
Measurementupstream v1.5.0, BF16 (fast)basal-rs, f16Unit
Single decision, median90.763.4–64.0ms (lower is better)
One question over HTTP, median (p50)11872ms (lower is better)
32 concurrent clients6.821.8requests/s (higher is better)
Mixed load, 32 clients17119requests/min (higher is better)
16k-token document, 5 questions112.25.3s (lower is better)

Upstream ran on the same GPU in its fastest mode (fast). GPU power limit: 250 W, or 300 W for the 16k-token document. A token is a piece of text: the model reads input as a sequence of tokens.

All benchmarks Three models per GPU, p50 and p99 response times (median and the time within which 99% of responses complete), mixed load, long documents and questions with many options.
Architecture

How does it work?

basal-rs receives an HTTP request, builds prompts following the Basal protocol, computes them on a GPU or Mac, and returns a probability distribution over the options. The number's color indicates whose code handles that step.

basal-rs engine Basal protocol and models · Remek Kinas
01 Six stages of a request
Request reception, validation and planning, queue, batch, GPU forward pass, decision and response.
02 The model author's protocol
Each question is sent to the model twice, with options in two orders. The result is averaged and calibrated. basal-rs reproduces Remek Kinas's protocol without changes.
03 Shared document computed once
Questions about the same document share a prompt prefix, so the engine computes it only once.
04 Short requests do not wait behind long ones
Requests over 4096 tokens enter a separate lane, so short questions do not wait behind long documents.
05 The same response regardless of traffic
A question's result does not depend on which other requests share its computation.
06 CUDA and Metal, without Python
Custom kernels on NVIDIA GPUs and Apple Silicon Macs, using f16 by default.
Questions

Frequently asked questions

Yes. The engine uses the Apache 2.0 license, which also permits commercial use. The Basal model author released the models under Apache 2.0 as well. No account or API key is required. There are no limits.

The Basal models and their author, Remek Kinas. basal-rs reproduces his decision protocol: the same prompt, token IDs, both option orders and calibration. Our measurements describe engine performance and compatibility with upstream reference calculations. They do not evaluate model accuracy.

f16 stores numbers in 16 bits; FP32 uses 32 bits (full precision). Our reference is the FP32 computation used by the author to evaluate the models. basal-rs defaults to f16 with f32 accumulation, matching FP32 decisions on 899–900 of 900 questions. The only differences occur in near ties, such as 0.503 versus 0.497 in FP32. Upstream's fast mode uses BF16, another 16-bit format, and matches on 890–896 of 900 questions. We compare performance against that mode: 16 bits against 16 bits. For FP32 computation, set dtype: f32.

An Apple Silicon Mac, or an x86_64 Linux computer with an NVIDIA GPU with compute capability 8.0 or higher (A100, RTX 30xx or newer). GPU memory must hold model weights plus activations (intermediate results): basal-1.5-mini ~3.2 GB, 4.5B ~9.5 GB, max ~23 GB.

Yes. POST /v1/systemone follows the TypeSafe System One contract, including its error statuses. POST /v1/basal follows the upstream v1.5.0 server conventions. The application does not need to know that basal-rs is on the other end.

The decision protocol is the same; the computation differs. basal-rs arranges prompts into a prefix tree, computing a shared state only once. It has custom CUDA and Metal GPU code and a scheduler that prioritizes short requests over long ones. No Python is needed at runtime.

The server has no authentication and does not check who sends requests. By default, it accepts only local connections (127.0.0.1). The 0.0.0.0 address accepts network connections and is also used in the Docker image. It is intended for a trusted network or for use behind an authenticating proxy.

Roadmap

What works and what comes next

Version 1.0.0 aims to refine the single server: supported systems, GPUs, performance and integrations. Later, the engine is planned to support multiple GPUs and machines. Releases and changelogs are published on GitHub. Issues and pull requests are welcome for every item.

Released (0.1.x)

  • NVIDIA GPUs (CUDA) and Apple Silicon (Metal) GPU code for NVIDIA compute capabilities 8.0, 8.9 and 9.0 in one binary, plus Apple Silicon support.
  • basal-1.5 and basal-1.0 models basal-1.5 mini, 4.5B and max with multi, act, facts and evidence extensions, plus basal-1.0-4.5B.
  • One-command installation Homebrew, install.sh script, Docker image, basal doctor and basal update.
  • Prometheus metrics Request counts, response times, queues and batches, per model.

1.0.0

  • All Basal models Versions 1.0, 1.5, 1.8 and 2.0. Support for basal-1.0-1.5B will be added. The author has announced basal-1.8 and 2.0; support will follow publication of their weights.
  • Windows Server and command-line tool (CLI) on Windows with NVIDIA GPUs, alongside Linux and macOS.
  • Bash automation client basal client: questions from the command line, data from stdin or JSONL files, concurrent requests and a local model without starting a server. Usage examples
  • Performance Lower latency where upstream is currently faster: single decisions and states up to about 2k tokens on H100.
  • More GPUs GPU code and measurements for additional NVIDIA generations beyond the currently supported compute capabilities 8.0, 8.9 and 9.0.
  • More tests Broader decision-compatibility tests against the reference and server load tests.
  • API clients and SDKs Libraries for calling the server from application code without manually constructing HTTP requests.

After 1.0.0

  • GPU clusters Multiple GPUs share a request queue and a set of models.
  • Distributed computation Servers connect and distribute requests among themselves.
  • Cluster and deployment management Add machines and deploy new model versions from one place.
  • Admin dashboard Machine, model and queue status in the browser.
  • Encrypted connections (mTLS) Encrypted, mutually authenticated connections between clients, servers and cluster nodes.
Responsibilities

Who is responsible for what

basal-rs runs models but does not create them. Basal models make the decisions; responsibility for decision quality lies with the model author, Remek Kinas. The basal-rs engine is responsible for computation speed and for how the computation is carried out.

owner You

Your application

Sends a state (text or data to evaluate) and questions about it over HTTP. Receives decisions with probabilities.

owner IT SOL

basal-rs

The Rust engine and server that run the models. We are responsible for how quickly and predictably decisions are computed.

  • GPU computation with CUDA (NVIDIA) and Metal (Apple Silicon)
  • shared prompt prefixes computed only once (prefix tree)
  • request scheduling and GPU batching
  • TypeSafe System One API and an endpoint compatible with the upstream server
  • installer, Docker image and metrics
  • public performance and decision-compatibility measurements
owner Remek Kinas

Basal

Decision models. They make the decisions and determine the quality of the answers.

  • basal-1.5 models: mini, 4.5B and max
  • model training and weights (learned parameters)
  • decision protocol: prompts, reading option-letter scores, options in two orders
  • probability calibration
  • the reference basal-serve server in Python, referred to here as upstream
owner SpeakLeash

Bielik v3.0

Polish base models (1.5B, 4.5B, 11B Instruct). Basal models are fine-tuned versions of these models.

Contact

Get in touch

basal-rs is an open project, and contributions are welcome. Send feedback, ideas, bug reports and pull requests in Polish or English.

  • Bugs and ideas

    Differences from upstream, performance issues and new GPU support. Use the GitHub issue form; Polish is welcome too.

    Open an issue
  • Pull requests

    Code and documentation fixes, new GPUs and performance work. For larger changes, open an issue first to agree on the approach and measurements.

    How to contribute
  • Questions

    Choosing a model, GPU or server configuration.

    GitHub Discussions
  • Security vulnerabilities

    Please do not report these publicly. Report privately through GitHub Security Advisories or to dev@itsol.tech.

    Report privately