vLLM
Signed model services behind the Router, offline at runtime.
Serving todayPrivate AI infrastructure
TensorBreeze turns the GPU machines you already own — Windows or Linux — into one OpenAI-compatible endpoint on your own network. No port forwarding. No Kubernetes.
routed → RTX 5090 · windows-tower · WSL2
Three settings. No TensorBreeze-specific code.
One fleet, many machines
Applications talk to one endpoint. Breeze Router matches each request to a healthy GPU with the memory to serve it, and keeps applications insulated from worker credentials and network topology. The Router binds to loopback beside your applications — nothing on your network can reach it — and requests travel to workers over node-initiated tunnels.
your app→Breeze Router · loopback→Breeze Controller→GPU worker · outbound tunnel
Enroll each machine with a one-time claim. Workers connect outward over HTTPS — nothing listens, nothing is port-forwarded.
One OpenAI-compatible endpoint with scoped per-app keys and model aliases. Requests land on the GPU that can actually serve them.
Models run in narrow, signed service profiles on the machine with the VRAM — your gaming PC included.
Breeze Console
Live utilization, memory, temperature, and process telemetry for every enrolled machine — alongside queueing, placement, and an append-only activity trail. These are real Console views of a demo fleet. Flip through them.
First-class Windows workers
Most orchestration stacks stop at Linux. TensorBreeze treats a Windows machine with WSL2 as a first-class worker — enrolled, monitored, and serving models like any other node, whether it’s a workstation or the PC you game on.
Planned Apple Silicon is next: a macOS worker is on the roadmap, so the GPU in your Mac can join the same fleet.
Running in the real world
TensorBreeze’s first production workload is its own proof: an application on its stable host routes chat completions to a 32-billion-parameter model served by vLLM on an RTX 5090 — inside WSL2, on a different machine, behind no open inbound ports. The application changed three settings.
Designed for private infrastructure
Workers connect outward and expose narrow inference contracts — never SSH, filesystems, or arbitrary commands. Every capability has its own credential, route, and schema, and everything dangerous ships switched off.
What it’s for
Applications stay on the host you trust; inference goes wherever the VRAM is. Three shapes this takes:
No magic, on purpose
Clear boundaries are a feature. These are permanent product truths, not temporary limitations.
Apps
Breeze separates client apps from inference providers, giving each integration only the credential and network access it needs. The catalog grows deliberately — reviewed and scoped, not scraped.
Signed model services behind the Router, offline at runtime.
Serving todayPrivate chat through a scoped Breeze connector.
Early accessYour own tools and agents point at the Router like any other endpoint.
By designEarly access